BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143045Z
LOCATION:B312-B313A
DTSTART;TZID=America/New_York:20241121T153000
DTEND;TZID=America/New_York:20241121T170000
UID:submissions.supercomputing.org_SC24_sess372@linklings.com
SUMMARY:Scale–Out Interconnects
DESCRIPTION:Network-Offloaded Bandwidth-Optimal Broadcast and Allgather fo
 r Distributed AI\n\nIn the Fully Sharded Data Parallel (FSDP) training pip
 eline, collective operations can be interleaved to maximize the communicat
 ion/computation overlap. In this scenario, outstanding operations such as 
 Allgather and Reduce-Scatter can compete for the injection bandwidth and c
 reate pipeline bubbles. ...\n\n\nMikhail Khalilov (ETH Zürich), Salvatore 
 Di Girolamo (NVIDIA Corporation), Marcin Chrapek (ETH Zürich), Rami Nudelm
 an and Gil Bloch (NVIDIA Corporation), and Torsten Hoefler (ETH Zürich)\n-
 --------------------\nUNR: Unified Notifiable RMA Library for HPC\n\nRemot
 e Memory Access (RMA) enables direct access to remote memory to achieve hi
 gh performance for HPC applications. However, most modern parallel program
 ming models lack schemes for the remote process to detect the completion o
 f RMA operations. Many previous works have proposed programming models an.
 ..\n\n\nGuangnan Feng and Jiabin Xie (Sun Yat-sen University, Guangzhou, C
 hina); Dezun Dong (National Key Laboratory of Parallel and Distributed Com
 puting); and Yutong Lu (Sun Yat-sen University, Guangzhou, China)\n-------
 --------------\nhZCCL: Accelerating Collective Communication with Co-Desig
 ned Homomorphic Compression\n\nAs network bandwidth lags behind increasing
  computing power, efficient collective communication is a major challenge 
 for exascale applications. Traditional approaches use error-bounded lossy 
 compression to accelerate collective operations but suffer from the costly
  decompression-operation-compressio...\n\n\nJiajun Huang (University of Ca
 lifornia, Riverside); Sheng Di (Argonne National Laboratory (ANL)); Xiaodo
 ng Yu (Stevens Institute of Technology); Yujia Zhai, Jinyang Liu, and Zizh
 e Jian (University of California, Riverside); Xin Liang (University of Ken
 tucky); Kai Zhao (Florida State University); Xiaoyi Lu (University of Cali
 fornia, Merced); Zizhong Chen (University of California, Riverside); and F
 ranck Cappello, Yanfei Guo, and Rajeev Thakur (Argonne National Laboratory
  (ANL))\n\nTag: Data Compression, Data Movement and Memory, Distributed Co
 mputing, Message Passing, Network\n\nRegistration Category: Tech Program R
 eg Pass\n\nSession Chair: Daniele De Sensi (Sapienza University of Rome)
END:VEVENT
END:VCALENDAR
