BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143140Z
LOCATION:B303
DTSTART;TZID=America/New_York:20241118T093000
DTEND;TZID=America/New_York:20241118T100000
UID:submissions.supercomputing.org_SC24_sess748_ws_pmbsf121@linklings.com
SUMMARY:Comprehensive Performance Modeling and System Design Insights for 
 Foundation Models
DESCRIPTION:Shashank Subramanian, Ermal Rrapaj, and Peter Harrington (Lawr
 ence Berkeley National Laboratory (LBNL)); Smeet Chheda (Stony Brook Unive
 rsity); and Steven Farrell, Brian Austin, Samuel Williams, Nicholas Wright
 , and Wahid Bhimji (Lawrence Berkeley National Laboratory (LBNL))\n\nGener
 ative AI, in particular large transformer models, are increasingly driving
  HPC system design in science and industry. We analyze performance charact
 eristics of such transformer models and discuss their sensitivity to the t
 ransformer type, parallelization strategy, and HPC system features (accele
 rators and interconnects). We utilize a performance model that allows us t
 o explore this complex design space and highlight its key components. We f
 ind that different transformer types demand different parallelism and syst
 em characteristics at different training regimes. Large language models ar
 e performant with 3D parallelism and amplify network needs only at pre-tra
 ining scales with reduced dependence on accelerator capacity and bandwidth
 . On the other hand, long-sequence transformers, representative of scienti
 fic foundation models, place a more uniform dependence on network and capa
 city with necessary 4D parallelism. Our analysis emphasizes the need for c
 loser performance modeling of different transformer types, keeping system 
 features in mind, and demonstrates a path towards this.\n\nTag: Accelerato
 rs, Modeling and Simulation, Performance Evaluation and/or Optimization To
 ols\n\nRegistration Category: Workshop Reg Pass\n\nSession Chairs: Simon H
 ammond (National Nuclear Security Administration (NNSA)); Sascha Hunold (T
 echnical University of Vienna); and Steven A. Wright (University of York, 
 England)\n\n
END:VEVENT
END:VCALENDAR
