The COVID-19 High Performance Computing Consortium

Bringing together the Federal government, industry, and academic leaders to provide access to the world’s most powerful high-performance computing resources in support of COVID-19 research.

  • 42
    Members
  • 6.4m
    Cores
  • 114
    Projects
  • 49.8k
    GPUs
  • 141.4k
    Nodes
  • 603
    Petaflops

The COVID-19 High Performance Computing Consortium

The COVID-19 High Performance Computing (HPC) Consortium is a unique private-public effort spearheaded by the White House Office of Science and Technology Policy, the U.S. Department of Energy and IBM to bring together federal government, industry, and academic leaders who are volunteering free compute time and resources on their world-class machines.

Read about the available resources below.

International Government Agencies and National Laboratories

  • Korea Institute of Science and Technology Information (KISTI)
  • Ministry of Education, Culture, Sports, Science and Technology (MEXT)-JAPAN
  • RIKEN Center for Computational Science (R-CCS)

Department of Energy National Laboratories

  • Argonne National Laboratory
  • Lawrence Livermore National Laboratory
  • Los Alamos National Laboratory
  • Oak Ridge National Laboratory
  • Lawrence Berkeley National Laboratory
  • Sandia National Laboratories
  • Idaho National Laboratory

Industry

  • IBM
  • Amazon Web Services
  • AMD
  • D. E. Shaw Research
  • Dell Technologies
  • Google Cloud
  • Hewlett Packard Enterprise
  • Microsoft
  • NVIDIA
  • Intel

Academia

  • Massachusetts Institute of Technology
  • Rensselaer Polytechnic Institute
  • University of Illinois
  • University of Texas at Austin
  • University of California - San Diego
  • Carnegie Mellon University
  • University of Pittsburgh
  • Indiana University
  • Massachusetts Green High Performance Computing Center (MGHPCC)
  • University of Wisconsin-Madison
  • Ohio Supercomputer Center
  • UK Digital Research Infrastructure
  • CSCS – Swiss National Supercomputing Centre
  • SNIC PDC – Swedish National Infrastructure for Computing, Center for High Performance Computing
  • Arizona State University
  • University of Alabama-Birmingham
  • Purdue University

Federal Agencies

  • National Science Foundation
    • XSEDE
    • Pittsburgh Supercomputing Center (PSC)
    • Texas Advanced Computing Center (TACC)
    • San Diego Supercomputer Center (SDSC)
    • National Center for Supercomputing Applications (NCSA)
    • Indiana University Pervasive Technology Institute (IUPTI)
    • Open Science Grid (OSG)
    • Purdue University Research Computing (RCAC)
  • NASA

Consortium collaborating initiatives represent efforts organized regionally around the world that are also working to accelerate research for fighting COVID-19. The Consortium is working with these initiatives and collaborating to share knowledge gained from our respective efforts.

  • EU PRACE COVID-19 Initiative
  • NCI Australia and Pawsey Supercomputing Centre

Consortium affiliates provide a range of computing services and expertise that can enhance and accelerate the research for fighting COVID-19. Matched proposals will have access to resources and help from Consortium affiliates, provided for free, enabling rapid and efficient execution of complex computational research programs.

  • Atrio
  • Data Expedition, Inc.
  • Flatiron
  • Fluid Numerics
  • Immortal Hyperscale InterPlanetary Fabrics
  • MathWorks
  • The HDF Group
  • Raptor Computing Systems, LLC
  • SAS

Project Proposal

Please note, as of May 1, 2022, the COVID-19 HPC Consortium is no longer accepting requests for allocations of resources and services to support the pandemic response. Those with active allocations will still have access to their allocation through the previously stated period.

To all those whose research projects have been supported by the Consortium, thank you for your research and impactful results.

For those needing access to additional resources, you can either work directly with providers you might already be working with, or pursue allocations through available programs and other opportunities.

HPC Resources

Consortium members and affiliates manage a range of computing capabilities: from small clusters to some of the largest supercomputers in the world. They offer not only computational resources, but also software, services, and deep technical expertise to help COVID-19 researchers execute complex computational research programs. Browse available member and affiliate resources here.

  • 603
    Petaflops
  • 49k
    GPUs

Members

IBM Cloud

Compute:

bx2-16x64 with 16 vCPUs, 64GB RAM and 32 Gbps
bx2-48x192 with 48 vCPUs, 192 GB RAM and 80 Gbps

cx2-16x32 with 16 vCPUs, 32 GB RAM and 32 Gbps
cx2-32x64 with 32 vCPUs, 64 GB RAM and 64 Gbps

mx2-16x128 with 16 vCPUs, 128 GB RAM and 32 Gbps
mx2-32x256 with 32 vCPUs, 256 GB RAM and 64 Gbps

Storage:

Cloud Object Storage

Job Scheduling:

IBM Spectrum LSF managed instances for easy job submission and monitoring.

IBM Research

IBM Research is providing our WSC 2.8 PF, 54 node, IBM POWER9/NVIDIA Volta high performance computing cluster as well as software tools to help accelerate discovery.

2 x POWER9 CPU per node, 22 cores per CPU
6 x NVIDIA Volta GPUs per node (336 total)
512 GiB DRAM per node
1.4 TB NVMe per node
2 x Mellanox EDR InfiniBand
2 PB IBM Spectrum Scale storage

Deep Search: The Deep Search platform for COVID-19 helps researchers to quickly find and aggregate information in the exponentially growing literature related to COVID-19. Examples of such information are the list of all reported used drugs so far.
Drug Candidate Exploration: To help researchers generate potential new drug candidates for COVID-19, we have applied our novel AI generative frameworks to three COVID-19 targets and have generated 3000 novel molecules. We are sharing these molecules under a Creative Commons license.
Functional Genomics Platform: IBM Functional Genomics Platform is a cloud-based data repository that accelerates the study of microbial life at scale with specifically curated molecular sequence data to fight COVID-19.

Beskow, a 2.5 PFLOPS Cray XC40, 2,060 nodes with dual Intel Haswell (16 cores/socket) and Broadwell (18 cores/socket) CPUs.

The Korea Institute of Science and Technology Information (KISTI) serves as a national supercomputing center of Korea, providing supercomputing and high-performance research networking facilities to Korean researchers. Together, with our hope to help accelerate the understanding of the COVID-19 virus for development of treatments and vaccines, KISTI is intended to contribute to the HPC Consortium of COVID-19 by offering access to the KISTI-5 supercomputer called Nurion. KISTI has also had a great experience in participating to the European project called WISDOM, a grid-enabled drug discovery initiative against malaria, several years ago in the era of Grid computing, where two teams in Korea joined the WISDOM collaboration by (1) offering computing resources along with relevant technology and (2) in-vitro testing to the initiative, respectively. KISTI still maintains the technology (http://htcaas.kisti.re.kr/wiki) to facilitate the conducting of large-scale virtual screening experiments to identify small molecule drug candidates on top of multiple computing platforms.

Nurion | 25.7PF, 8437 nodes, Cray CS500
[8305 nodes] Intel Xeon Phi 7250 (KNL) 68C 1.4GHz + [132 nodes] 2 x Intel Xeon 6148 (Skylake) 20C 2.4GHz
96GB DDR4 and 16 GB MCDRAM memory per KNL node
192GB DDR4 memory per Skylate node
0.8PB DDN IME flash storage (burst buffer)
20PB Lustre Filesystem
10PB IBM TS 4500 Tape Storage
Intel Omni-Path, Fat-Tree, 50% Blocking

2.2 PF, 7375 nodes, IBM POWER8/9, Intel Xeon

DOE/NNSA LLNL Lassen | 23 PF, 788 compute nodes, IBM POWER9/NVIDIA Volta GV100
DOE/NNSA LLNL Quartz | 3.2 PF, 3004 compute nodes, Intel Xeon Broadwell
DOE/NNSA LLNL Pascal | 0.9 PF, 163 compute nodes, Intel Xeon Broadwell CPUs/NVIDIA Pascal P100
DOE/NNSA LLNL Ray | 1.0 PF, 54 compute nodes, IBM POWER8/NVIDIA Pascal P100
DOE/NNSA LLNL Surface | 506 TF, 158 compute nodes, Intel Xeon Sandy Bridge/NVIDIA Kepler K40m
DOE/NNSA LLNL Syrah | 108 TF, 316 compute nodes, Intel Xeon Sandy Bridge
DOE/NNSA LANL Grizzly | 1.8 PF, 1490 compute nodes, Intel Xeon Broadwell
DOE/NNSA LANL Snow | 445 TF, 368 compute nodes, Intel Xeon Broadwell
DOE/NNSA LANL Badger | 790 TF, 660 compute nodes, Intel Xeon Broadwell
DOE/NNSA SNL Solo | 460 TF, 374 compute nodes, Intel Xeon Broadwell

The supercomputer Fugaku is Japan’s flagship supercomputer, developed mainly via collaboration between RIKEN R-CCS and Fujitsu, and to be commissioned for operation in 2021; however, portions of its resources (approximately 89 PF) is being deployed a year in advance to combat COVID-19. The technical specifications of Fugaku are as follows:

Processor core ISA: Arm (Aarch64 v8 + 512 bit SVE)
48 + 2 or 4 cores per CPU chip, one CPU chip per node, ~400Gbps Tofu-D interconnect.
Total # Nodes: 158,976 nodes
Theoretical Peak Compute Performances: Boost Mode (CPU Frequency 2.2GHz)
64 bit Double Precision FP: 537 Petaflops
32 bit Single Precision FP: 1.07 Exaflops
16 bit Half Precision FP (AI training): 2.15 Exaflops
8 bit Integer (AI Inference): 4.30 Exaops
Theoretical Peak Memory Bandwidth: 163 Petabytes/s
Approximately 150 Petabytes of Lustre storage
System software: Red Hat Enterprise Linux, all standard programming languages, optimized numerical libraries, MPI, OpenMP, TensorFlow/PyTorch, etc.

For details refer to:
https://postk-web.r-ccs.riken.jp/spec.html
https://www.fujitsu.com/global/about/innovation/fugaku/specifications/

Transform research data into valuable insights and conduct large-scale analyses with the power of Google Cloud. As part of the COVID-19 HPC Consortium, Google is providing access to Google Cloud HPC resources for academic researchers.

N2D Instances | Up to 224 vCPUs, High Memory Bandwidth, AMD EPYC ROME
· 896GB RAM
N1 Instances | Up to 416 vCPUs, General Purpose, Intel Skylake, Intel Ivy Bridge
· 12TB DRAM
· Up to 9TB Local SSD, 7TB Intel Optane
C2 Instances | Up to 60 vCPUs , HPC Optimized, Intel Cascade Lake
· 240GB DRAM
GPUs | NVIDIA Tesla K80, Tesla P100, Tesla P4, Tesla V100, Tesla T4
Tensor Processing Units V3 | 420 teraflops, 128 GB memory
Storage:
· Google Cloud Filestore
· Netapp Cloud Volumes
· DDN Exascaler Managed Lustre
· Open-source Lustre
· Google Cloud Storage

Intel will provide HPC /AI and HLS subject matter experts and engineers to collaborate on COVID-19 code enhancements to benefit the community. Intel will also provide licenses for High Performance Computing software development tools for the research programs selected by the COVID-19 HPC Consortium. The integrated tool suites include Intel's C++ and Fortran Compilers, performance libraries, and performance-analysis tools.

A task force of NVIDIA researchers and data scientists with expertise in AI and HPC will help optimize research projects on the Consortium’s supercomputers. The NVIDIA team has expertise across a variety of domains, including AI, supercomputing, drug discovery, molecular dynamics, genomics, medical imaging and data analytics. NVIDIA will also contribute the packaging of software for relevant AI and life-sciences software applications through NVIDIA NGC, a hub for GPU-accelerated software. The company is also providing compute time on an AI supercomputer, SaturnV.

11.1 PF, 252 nodes POWER9/Volta

2 x IBM POWER9 CPU per node, 20 cores per CPU
6 x NVIDIA Tesla GV100 per node
32 GB HBM per GPU
512 GB DRAM per node
1.6 TB NVMe per node
Mellanox EDR InfiniBand
11 PB IBM Spectrum Scale storage

11.69 PF, 4292 nodes, Intel KNL

1 x Intel KNL 7230 per node, 64 cores per CPU
192 GB DDR4, 16GB MCDRAM memory per node
128 GB local storage per node
Aries dragonfly network
10 PB Lustre + 1 PB IBM Spectrum Scale storage

Full details available at: https://www.alcf.anl.gov/alcf-resources

By expanding our existing AI for Health program, Microsoft will give access to our Azure cloud and High-Performance Computing capabilities. Our team of AI for Health data science experts, whose mission is to improve the health of people and communities worldwide, is also open to collaborations with COVID-19 researchers as they tackle this critical challenge. More broadly, Microsoft’s research scientists across the world, spanning computer science, biology, medicine, and public health, will be available to provide advice and collaborate per mutual interest.

AI for Health web site: https://www.microsoft.com/en-us/ai/ai-for-health
Azure web site: https://azure.microsoft.com/en-us/
Azure HPC web site: https://azure.microsoft.com/en-us/solutions/high-performance-computing/

Compute:
HBv2: 120 cores AMD Rome, 480 GB Memory, 1 TB NVMe, per node. Mellanox HDR InfiniBand.
HB: 60 cores AMD Naples, 240 GB Memory, 700 GB NVMe, per node. Mellanox EDR InfiniBand.
HC: 44 cores Intel Skylake,352 GB Memory, 700 GB NVMe, per node. Mellanox EDR InfiniBand
NDv2: 8 x Nvidia V100 with NVLINK, 40 cores Intel Skylake, 672GB Memory, 3 TB SSD, per node. Mellanox EDR InfiniBand
NCv3: 8 x Nvidia V100, 24 vCPUs Intel Broadwell, 448 GB Memory, 3 TB SSD, per node. Mellanox FDR InfiniBand

Storage:
Azure HPC Cache
Azure NetApp Files
Azure Blob Storage
Cray ClusterStor

Cluster Services:
Azure CycleCloud: LSF, Slurm, PBSPro, HTCondor, Grid Engine cluster orchestration

19.39 PF, 17609 nodes Intel Xeon

AITKEN | 3.69 PF, 1,152 nodes, Intel Xeon
ELECTRA | 8.32 PF, 3,456 nodes, Intel Xeon
PLEIDES | 7.09 PF, 11,207 nodes, Intel Xeon, NVIDIA K40, Volta GPUs
ENDEAVOR | 32 TF, 2 nodes, Intel Xeon
MEROPE | 253 TF, 1792 nodes, Intel Xeon

200 PF, 4608 nodes, IBM POWER9/NVIDIA Volta

2 x IBM POWER9 per node
42 TF per node
6 x NVIDIA Volta GPUs per node
512 GB DDR4 + 96 GB HBM2 (GPU memory) per node
1600 GB per node
2 x Mellanox EDR IB adapters (100Gbps per adapter)
250 PB, 2.5 TB/s, IBM Spectrum Scale storage

OSC Owens | 1.6 PF, 824 nodes Intel Xeon/Pascal
2 x Intel Xeon (28 cores per node, 48 cores per big-mem node)
160 NIVIDIA P100 GPUs (1 per node)
128 GB per node (1.5TB per big-mem node)
Mellanox EDR Infiniband
12.5 PB Project and Scratch storage

OSC Pitzer | 1.3.PF, 260 nodes Intel Xeon/Volta
2 x Intel Xeon (40 cores per node, 80 cores per big-mem node)
64 NIVIDIA V100 GPUs (2 per node)
192 GB per node (384 GB per GPU node, 3TB per big-mem node)
Mellanox EDR Infiniband
12.5 PB Project and Scratch storage

Affiliates

Atrio will assist researchers studying COVID-19 in creating and optimizing performance of application containers (e.g. CryoEM processing application suite), as well as performance-optimized deployment of those application containers on to any of HPC Consortium members' computational platforms and specifically onto high performing GPU and CPU resources. Our proposal is two fold - one is additional computational resources, and another, equally important, is support for COVID-19 researchers with an easy way to access and use HPC Consortium computational resources. That support consists of creating application containers for researchers, optimizing their performance, and an optional multi-site container and cluster management software toolset.

Immortal is providing licensed access to components of its platform to organizations which are (a) investigating the nature of the COVID-19 virus, (b) developing products for therapeutic breakthroughs, and (c) conducting R&D to build COVID-19 vaccines. Immortal's platform aggregates and orchestrates large magnitudes of applications, services, data, and resources across multiple Clouds, multiple supercomputers, or a combination. The platform is suitable for organizations operating on problems which need computation and data management at scales of hundreds of petaflops and hundreds of petabytes.

MathWorks will help researchers studying COVID-19 to scale their parallel MATLAB algorithms to the cloud and to HPC resources provided by this HPC Consortium. Our offering includes:
•\tFree access to MATLAB® and Simulink® on the allocated computing resources.
•\tSupport for parallelizing and scaling researcher’s algorithms.
As an example, see Ventilator research at Duke University.

The HDF Group helps scientists use open source HDF5 effectively, including offering general usage and performance tuning advice, and helping to troubleshoot any issues that arise. Our engineers will be available to assist you in applying HPC and HDF® technologies together for your Covid-19 research.