Publications
2025
TRACER Software Switching at Waveform Timescales
2025 IEEE High Performance Extreme Computing Conference (HPEC), 1-7, 2025
FedPaI: Achieving Extreme Sparsity in Federated Learning via Pruning at Initialization
arXiv preprint arXiv:2504.00308, 2025
TRACER, the Next Generation Ultra-wideband Spectrum Processor
MILCOM 2025-2025 IEEE Military Communications Conference (MILCOM), 396-403, 2025
2024
Timely Wildfire Perimeter Mapping for Unmanned Aerial Platforms
Applied Industrial Spectroscopy, FD1. 7, 2024
Evaluating Deep Learning Recommendation Model Training Scalability with the Dynamic Opera Network
Proceedings of the 4th Workshop on Machine Learning and Systems, 169-175, 2024
MoQ: Mixture-of-format Activation Quantization for Communication-efficient AI Inference System
NeurIPS 2024 Workshop Machine Learning with new Compute Paradigms, 2024
2023
Distributed edge machine learning pipeline scheduling with reverse auctions
2023 Eighth International Conference on Fog and Mobile Edge Computing (FMEC …, 2023
Quantpipe: Applying adaptive post-training quantization for distributed transformer pipelines in dynamic edge environments
ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and …, 2023
2022
Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices
2022 25th Euromicro Conference on Digital System Design (DSD), 298-307, 2022
2021
Distributed and Heterogeneous SAR Backprojection with Halide
2021 IEEE High Performance Extreme Computing Conference (HPEC), 1-9, 2021
Demonstration of a fully neural network based synthetic aperture radar processing pipeline for image formation and analysis
Sensors, Systems, and Next-Generation Satellites XXV 11858, 98-109, 2021
Pipeline parallelism for inference on heterogeneous edge computing
arXiv preprint arXiv:2110.14895, 2021
2020
A case study and characterization of a many-socket, multi-tier numa hpc platform
2020 IEEE/ACM 6th Workshop on the LLVM Compiler Infrastructure in HPC (LLVM …, 2020
Compiler abstractions and runtime for extreme-scale sar and cfd workloads
2020 IEEE/ACM Fifth International Workshop on Extreme Scale Programming …, 2020
RDANet: a deep learning based approach for synthetic aperture radar image formation
arXiv preprint arXiv:2001.08202, 2020
2019
Increased Fault-Tolerance and Real-Time Performance Resiliency for Stream Processing Workloads through Redundancy
2019 IEEE International Conference on Services Computing (SCC), 51-55, 2019
Computational requirements for real-time ptychographic image reconstruction
Applied Optics 58 (7), B19-B27, 2019
2018
Pacer: Automated Feedback-Based Vertical Elasticity for Heterogeneous Soft Real-Time Workloads
2018 IEEE/ACM 11th International Conference on Utility and Cloud Computing …, 2018
Reducing tail latencies while improving resiliency to timing errors for stream processing workloads
2018 IEEE/ACM 11th International Conference on Utility and Cloud Computing …, 2018
2017
A comparison of system performance on a private openstack cloud and amazon ec2
2017 IEEE 10th International Conference on Cloud Computing (CLOUD), 310-317, 2017
Dynamically Improving Resiliency to Timing Errors for Stream Processing Workloads
2017 18th International Conference on Parallel and Distributed Computing …, 2017
Load balancing for minimizing deadline misses and total runtime for connected car systems in fog computing
2017 IEEE International Symposium on Parallel and Distributed Processing …, 2017
2016
Hypervisor performance analysis for real-time workloads
2016 IEEE High Performance Extreme Computing Conference (HPEC), 1-7, 2016
Automated demand-based vertical elasticity for heterogeneous real-time workloads
2016 IEEE 9th International Conference on Cloud Computing (CLOUD), 831-834, 2016
2015
Supporting high performance molecular dynamics in virtualized clusters using IOMMU, SR-IOV, and GPUDirect
ACM SIGPLAN Notices 50 (7), 31-38, 2015
Heterogeneous cloud computing: The way forward
Computer 48 (01), 59-61, 2015
2014
GPU passthrough performance: A comparison of KVM, Xen, VMWare ESXi, and LXC for CUDA and OpenCL applications
2014 IEEE 7th international conference on cloud computing, 636-643, 2014
Bridging the virtualization performance gap for HPC using SR-IOV for InfiniBand
2014 IEEE 7th International Conference on Cloud Computing, 627-635, 2014
Evaluating GPU passthrough in Xen for high performance cloud computing
2014 IEEE international parallel & distributed processing symposium …, 2014
2013
A practical characterization of a NASA SpaceCube application through fault emulation and laser testing
2013 43rd Annual IEEE/IFIP International Conference on Dependable Systems …, 2013
2012
Silent data corruption and embedded processing with nasa's spacecube
IEEE Embedded Systems Letters 4 (2), 33-36, 2012
Applying radiation hardening by software to fast lossless compression prediction on FPGAs
2012 IEEE Aerospace Conference, 1-10, 2012
Integrating high performance file systems in a cloud computing environment
2012 SC Companion: High Performance Computing, Networking Storage and …, 2012
2011
Heterogeneous cloud computing
2011 IEEE International Conference on Cluster Computing, 378-385, 2011
Programming models and development software for a space-based many-core processor
2011 IEEE Fourth International Conference on Space Mission Challenges for …, 2011
The PowerPC 405 memory sentinel and injection system
2011 IEEE 19th Annual International Symposium on Field-Programmable Custom …, 2011
Software-based fault tolerance for the Maestro many-core processor
2011 Aerospace Conference, 1-12, 2011
Software fault tolerance methodology and testing for the embedded PowerPC
2011 Aerospace Conference, 1-9, 2011
Fftw and complex ambiguity function performance on the maestro processor
2011 Aerospace Conference, 1-8, 2011
Autonomous on-board processing for sensor systems: Initial fault tolerance and autonomy results
Earth Science Technology Forum, 2011
2010
Database Searching with Profile-Hidden Markov Models on Reconfigurable and Many-Core Architectures
Bioinformatics: high performance parallel computer architectures 2, 203, 2010
Computation Checkpointing and Migration
Nova Science Publishers, Inc., 2010
Applying graphics processor units to Monte Carlo dose calculation in radiation therapy
Journal of Medical Physics 35 (2), 120-122, 2010
2009
Improving MPI-HMMER's scalability with parallel I/O
Parallel & Distributed Processing. IPDPS 2009. IEEE International …, 2009
Evaluating the use of GPUs in liver image segmentation and HMMER database searches
2009 Ieee International Symposium on Parallel & Distributed Processing, 1-12, 2009
A fault-tolerant strategy for virtualized HPC clusters
The Journal of Supercomputing 50, 209-239, 2009
2008
Replication-based fault tolerance for MPI applications
IEEE Transactions on Parallel and Distributed Systems 20 (7), 997-1010, 2008
A comparison of virtualization technologies for HPC
22nd International Conference on Advanced Information Networking and …, 2008
Accelerating molecular dynamics simulations with gpus
ISCA PDCCS, 44-49, 2008
Optimized Cluster‐Enabled HMMER Searches
Grid computing for bioinformatics and computational biology, 51-70, 2008
Enabling interactive jobs in virtualized data centers
Cloud Computing and Its Applications, 2008
2007
MPI-HMMER-Boost: Distributed fpga acceleration
The Journal of VLSI Signal Processing 48 (3), 223-238, 2007
Failure Prediction and Scalable Checkpointing for Reliable Large-Scale Grid Computing
IEEE HPDC’07, 2007
Wireless sensor network security: A survey
Security in distributed, grid, mobile, and pervasive computing, 367, 2007
A comprehensive user-level checkpointing strategy for MPI applications
Technical report, TR 2007-1, 2007
Wireless sensor network security: A survey, in book chapter of security
in Distributed, Grid, and Pervasive Computing, Yang Xiao (Eds, 0-849, 2007
Fault-tolerant techniques for high performance computing and a bioinformatics application
ProQuest, 2007
A scalable asynchronous replication-based strategy for fault tolerant MPI applications
High Performance Computing, HiPC 2007, 257-268, 2007
2006
An adaptive heterogeneous software DSM
2006 International Conference on Parallel Processing Workshops (ICPPW'06), 8 …, 2006
Accelerating HMMer searches on Opteron processors with minimally invasive recoding
20th International Conference on Advanced Information Networking and …, 2006
Accelerating the HMMER sequence analysis suite using conventional processors
20th International Conference on Advanced Information Networking and …, 2006
Data Conversion for Heterogeneous Migration/Checkpointing
High-performance computing: paradigm and infrastructure 44, 241, 2006
Application-level checkpointing techniques for parallel programs
Distributed Computing and Internet Technology, 221-234, 2006
2003
Data conversion for process/thread migration and checkpointing
2003 International Conference on Parallel Processing. Proceedings …, 2003