BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143140Z
LOCATION:B309
DTSTART;TZID=America/New_York:20241119T103000
DTEND;TZID=America/New_York:20241119T110000
UID:submissions.supercomputing.org_SC24_sess380_pap547@linklings.com
SUMMARY:LLM-Pilot: Characterize and Optimize Performance of your LLM Infer
 ence Services
DESCRIPTION:Malgorzata Lazuka (IBM Zurich Research Laboratory, ETH Zürich)
  and Andreea Anghel and Thomas Parnell (IBM Zurich Research Laboratory)\n\
 nAs Large Language Models (LLMs) are rapidly growing in popularity, LLM in
 ference services must be able to serve requests from thousands of users wh
 ile satisfying performance requirements. The performance of an LLM inferen
 ce service is largely determined by the hardware onto which it is deployed
 , but understanding of which hardware will deliver on performance requirem
 ents remains challenging. In this work we present LLM-Pilot - a first-of-i
 ts-kind system for characterizing and predicting performance of LLM infere
 nce services. LLM-Pilot performs benchmarking of LLM inference services, u
 nder a realistic workload, across a variety of GPUs, and optimizes the ser
 vice configuration for each considered GPU to maximize performance. Finall
 y, using this characterization data, LLM-Pilot learns a predictive model, 
 which can be used to recommend the most cost-effective hardware for a prev
 iously unseen LLM. Compared to existing methods, LLM-Pilot can deliver on 
 performance requirements 33% more frequently, whilst reducing costs by 60%
  on average.\n\nTag: Algorithms, Artificial Intelligence/Machine Learning,
  Data Movement and Memory, Graph Algorithms\n\nRegistration Category: Tech
  Program Reg Pass\n\nSession Chair: Jay Lofstead (Sandia National Laborato
 ries, University of New Mexico)\n\n
END:VEVENT
END:VCALENDAR
