BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143141Z
LOCATION:B312-B313A
DTSTART;TZID=America/New_York:20241119T163000
DTEND;TZID=America/New_York:20241119T170000
UID:submissions.supercomputing.org_SC24_sess375_pap565@linklings.com
SUMMARY:Exploring GPU-to-GPU Communication: Insights into Supercomputer In
 terconnects
DESCRIPTION:Daniele De Sensi (Sapienza University of Rome); Lorenzo Pichet
 ti and Flavio Vella (University of Trento, Italy); Tiziano De Matteis and 
 Zebin Ren (Vrije Universiteit Amsterdam); Luigi Fusco (ETH Zürich); Matteo
  Turisini and Daniele Cesarini (CINECA); Kurt Lust (University of Antwerp,
  Belgium); Animesh Trivedi (IBM); Duncan Roweth (Hewlett Packard Enterpris
 e (HPE)); Filippo Spiga and Salvatore Di Girolamo (NVIDIA Corporation); an
 d Torsten Hoefler (ETH Zürich)\n\nMulti-GPU nodes are increasingly common 
 in the rapidly evolving landscape of exascale supercomputers. On these sys
 tems, GPUs on the same node are connected through dedicated networks, with
  bandwidths up to a few terabits per second. However, gauging performance 
 expectations and maximizing system efficiency is challenging due to differ
 ent technologies, design options, and software layers. This paper comprehe
 nsively characterizes three supercomputers — Alps, Leonardo, and LUMI — ea
 ch with a unique architecture and design. We focus on performance evaluati
 on of intra-node and inter-node interconnects on up to 4096 GPUs, using a 
 mix of intra-node and inter-node benchmarks. By analyzing its limitations 
 and opportunities, we aim to offer practical guidance to researchers, syst
 em architects, and software developers dealing with multi-GPU supercomputi
 ng. Our results show that there is untapped bandwidth, and there are still
  many opportunities for optimization, ranging from network to software opt
 imization.\n\nTag: Accelerators, HPC Infrastructure, Performance Evaluatio
 n and/or Optimization Tools, State of the Practice\n\nRegistration Categor
 y: Tech Program Reg Pass\n\nAward Finalist: Best Paper Finalist\n\nSession
  Chair: Zachary Tschirhart (Sustainment)\n\n
END:VEVENT
END:VCALENDAR
