BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143139Z
LOCATION:Exhibit Hall A3
DTSTART;TZID=America/New_York:20241120T093000
DTEND;TZID=America/New_York:20241120T100000
UID:submissions.supercomputing.org_SC24_sess501_awd103@linklings.com
SUMMARY:Immense-Scale Machine Learning: The Big, the Small, and the Not Ri
 ght at All (2024 IEEE-CS Seymour Cray Computer Engineering Award)
DESCRIPTION:Norm Jouppi (Google)\n\nWe start with key aspects of the decad
 e-long evolution of Google’s Tensor Processing Systems.  These systems hav
 e unique requirements arising from serving billions of daily users across 
 multiple products in production environments.  Training Large Language ML 
 Models (LLMs) can require 100,000 accelerators working together in synchro
 ny for months.  And LLM inference can require exaFLOP/OP speeds with respo
 nse times of under a second.  This has driven adoption of lower-precision 
 numerics across the industry.  Finally, since random errors from various s
 ources are likely to occur at this immense scale during months-long traini
 ng runs, fault tolerance is also a key feature of immense-scale ML systems
 .\n\nRegistration Category: Tech Program Reg Pass\n\nSession Chairs: Venka
 tesh Kannan (Irish Centre for High‑End Computing (ICHEC)) and Scott Pakin 
 (Los Alamos National Laboratory (LANL))\n\n
END:VEVENT
END:VCALENDAR
