Presenter Full Schedule · Contributors · Organizations · Search ProgramShuangyan YangUniversity of California, MercedPresentationsPostersLM-Offload: Performance Model-Guided Generative Inference of Large Language Models with Parallelism Control12pm - 5pm EST Tuesday, 19 November 2024 B302-B305 TP XO/EX