BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143140Z
LOCATION:B312-B313A
DTSTART;TZID=America/New_York:20241121T160000
DTEND;TZID=America/New_York:20241121T163000
UID:submissions.supercomputing.org_SC24_sess372_pap299@linklings.com
SUMMARY:hZCCL: Accelerating Collective Communication with Co-Designed Homo
 morphic Compression
DESCRIPTION:Jiajun Huang (University of California, Riverside); Sheng Di (
 Argonne National Laboratory (ANL)); Xiaodong Yu (Stevens Institute of Tech
 nology); Yujia Zhai, Jinyang Liu, and Zizhe Jian (University of California
 , Riverside); Xin Liang (University of Kentucky); Kai Zhao (Florida State 
 University); Xiaoyi Lu (University of California, Merced); Zizhong Chen (U
 niversity of California, Riverside); and Franck Cappello, Yanfei Guo, and 
 Rajeev Thakur (Argonne National Laboratory (ANL))\n\nAs network bandwidth 
 lags behind increasing computing power, efficient collective communication
  is a major challenge for exascale applications. Traditional approaches us
 e error-bounded lossy compression to accelerate collective operations but 
 suffer from the costly decompression-operation-compression (DOC) workflow.
  We propose hZCCL, the first homomorphic compression-communication co-desi
 gn enabling direct operations on compressed data, avoiding expensive DOC o
 verhead. Alongside the co-design framework, we introduce a lightweight, mu
 lti-core CPU-optimized compressor and a homomorphic compressor with a runt
 ime heuristic to select efficient compression pipelines dynamically. We ev
 aluate hZCCL with up to 512 nodes and across five application datasets. Th
 e experimental results demonstrate that our homomorphic compressor achieve
 s a CPU throughput of up to 379.08 GB/s, surpassing the conventional DOC w
 orkflow by up to 36.53X. Moreover, our hZCCL-accelerated collectives outpe
 rform two state-of-the-art baselines, delivering speedups of up to 2.12X a
 nd 6.77X compared to original MPI collectives in single-thread and multi-t
 hread modes, respectively, while maintaining data accuracy.\n\nTag: Data C
 ompression, Data Movement and Memory, Distributed Computing, Message Passi
 ng, Network\n\nRegistration Category: Tech Program Reg Pass\n\nSession Cha
 ir: Daniele De Sensi (Sapienza University of Rome)\n\n
END:VEVENT
END:VCALENDAR
