Publications
Running AlphaFold3 on Distributed High-Throughput Computing Infrastructure: Scaling Workloads and Enabling Ultra-Large Predictions
Abstract
AlphaFold 3 (AF3) enables atomic-resolution prediction of biomolecular complexes, driving rapidly growing demand across the life sciences. However, its ∼ 750,GB reference database has effectively confined production deployments to systems with shared parallel filesystems, creating a major barrier for scalability. Distributed high-throughput computing (dHTC) platforms offer vast, heterogeneous compute capacity, but fundamentally lack the shared data infrastructure assumed by AF3. We present a data-aware deployment of AF3 for dHTC, implemented on the Center for High Throughput Computing (CHTC) and the Open Science Pool (OSPool). The workflow is decomposed into a CPU-bound data pipeline that executes on nodes with locally staged, scheduler-advertised databases, and a GPU-bound inference pipeline that opportunistically scales across distributed resources. Using CUDA Unified Virtual Memory …
- Date
- 2026
- Authors
- Daniel A Morales, Brian Lin, Mats Rynge, Christina L Koch, Brian Bockelman, Miron Livny
- Book
- Proceedings of the Practice and Experience in Advanced Research Computing 2026: Resilient Roots+ Empowered Communities
- Pages
- 1-5