Presentation
Copper: Cooperative Caching Layer for Scalable Data Loading in Exascale Supercomputers
DescriptionJob initialization time of dynamic executables increases as HPC jobs launch on a larger number of nodes and processes. This is due to the processes flooding the storage system with a tremendous number of I/O requests for the same files, leading to significant performance degradation, causing the nodes to remain idle for an extended period of time waiting for I/O resources, wasting valuable CPU cycles. In order to remove the I/O bottleneck occurring at job initialization time, a data loader capable of reducing storage-side congestion is necessary. In this paper, we introduce Copper, a read-only cooperative caching layer aimed to enable scalable data loading on a large number of nodes. We evaluate our data loading solution on the Aurora supercomputer located at Argonne National Laboratory. Our experiments show that Copper is able to load data in near constant time when scaling from 32 to 8,300 nodes.

