BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143138Z
LOCATION:B206
DTSTART;TZID=America/New_York:20241122T094500
DTEND;TZID=America/New_York:20241122T100000
UID:submissions.supercomputing.org_SC24_sess769_ws_adugai117@linklings.com
SUMMARY:llm-recipes: A Framework for Seamless Integration and Efficient Co
 ntinual Pre-Training of Large Language Models
DESCRIPTION:Kazuki Fujii, Taishi Nakamura, and Rio Yokota (Tokyo Institute
  of Technology)\n\nLarge Language Models (LLMs) have advanced natural lang
 uage processing but present challenges in training due to the complexity a
 nd resource demands. A significant issue is the inconsistency in checkpoin
 t formats across pre-training libraries, complicating the use of pre-train
 ed weights for continued training. To address this, we introduce llm-recip
 es, an open-source framework that streamlines the continual pre-training p
 rocess by enabling direct use of Hugging Face Transformers checkpoints wit
 hout conversion. This framework supports multi-node distributed training u
 sing PyTorch Fully Sharded Data Parallel (FSDP), enhancing scalability for
  large-scale models. Unlike existing tools, llm-recipes offers broader sup
 port for various model architectures and flexible training configurations,
  making it an adaptable solution for researchers and developers. Our exper
 iments demonstrate its effective scalability, with high training throughpu
 t up to 64 GPUs, confirming its suitability for large-scale distributed tr
 aining of LLMs.\n\nTag: Artificial Intelligence/Machine Learning\n\nRegist
 ration Category: Workshop Reg Pass\n\nSession Chairs: Charlie Catlett (Arg
 onne National Laboratory (ANL), University of Chicago); Fabrizio Gagliardi
  (Barcelona Supercomputing Center (BSC), Association for Computing Machine
 ry (ACM)); Neeraj Kumar (Pacific Northwest National Laboratory (PNNL)); Sa
 toshi Matsuoka (RIKEN, Tokyo Institute of Technology); Irina Rish (Univers
 ity of Montreal, Canada; Mila – Quebec AI Institute); Rick Stevens (Argonn
 e National Laboratory (ANL), University of Chicago); and Valerie Taylor (A
 rgonne National Laboratory (ANL), University of Chicago)\n\n
END:VEVENT
END:VCALENDAR
