BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143138Z
LOCATION:B309
DTSTART;TZID=America/New_York:20241119T113000
DTEND;TZID=America/New_York:20241119T120000
UID:submissions.supercomputing.org_SC24_sess380_pap482@linklings.com
SUMMARY:Efficient Weighted Graph Matching on GPUs
DESCRIPTION:Michael Mandulak (Rensselaer Polytechnic Institute (RPI)); Say
 an Ghosh, S. M. Ferdous, and Mahantesh Halappanavar (Pacific Northwest Nat
 ional Laboratory (PNNL)); and George Slota (Rensselaer Polytechnic Institu
 te (RPI))\n\nWeighted matching identifies a maximal subset of edges in a g
 raph with no common vertices. As a prototypical graph problem, matching ha
 s numerous applications in multi-level graph algorithms and machine learni
 ng. However, challenges arise in developing efficient, parallel graph matc
 hing methods on contemporary GPGPU systems due to general graph processing
  complexities, such as irregular memory access patterns and load imbalance
 s. Furthermore, increasingly massive graph sizes exceed available GPU memo
 ry while data dependencies and synchronization costs in multi-GPU massive-
 graph processing challenge sustainable scalability.\n\nConsidering these c
 hallenges, we present efficient approximation algorithms for locally-domin
 ant matching and demonstrate performance improvements via batching and dis
 tribution across multiple NVIDIA A100/V100 GPUs of NVIDIA DGX dense-GPU pl
 atform. Our pointer-based matching method exhibits 2-45x performance impro
 vements compared to state-of-the-art single-GPU and multithreaded CPU matc
 hing implementations on real-world and synthetic graphs. We show competiti
 ve quality comparisons and analysis of GPU-data distribution for efficient
 , practical graph matching on GPUs.\n\nTag: Algorithms, Artificial Intelli
 gence/Machine Learning, Data Movement and Memory, Graph Algorithms\n\nRegi
 stration Category: Tech Program Reg Pass\n\nSession Chair: Jay Lofstead (S
 andia National Laboratories, University of New Mexico)\n\n
END:VEVENT
END:VCALENDAR
