BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143037Z
LOCATION:B303
DTSTART;TZID=America/New_York:20241118T090000
DTEND;TZID=America/New_York:20241118T173000
UID:submissions.supercomputing.org_SC24_sess748@linklings.com
SUMMARY:PMBS24: The 15th International Workshop on Performance Modeling, B
 enchmarking, and Simulation of High-Performance Computer Systems
DESCRIPTION:The PMBS24 workshop is concerned with the comparison of high-p
 erformance computing systems through performance modeling, benchmarking or
  through the use of tools such as simulators. We are particularly interest
 ed in research which reports the ability to measure and make tradeoffs in 
 software/hardware co-design to improve sustained application performance. 
 We are also keen to capture the assessment of future systems. The aim of t
 his workshop is to bring together researchers, from industry and academia,
  concerned with the qualitative and quantitative evaluation and modeling o
 f high-performance computing systems. Authors are invited to submit novel 
 research in all areas of performance modeling, benchmarking and simulation
 , and we welcome research that brings together current theory and practice
 . We recognize that the term 'performance' has broadened to include power 
 consumption and reliability, and that performance modeling is practiced th
 rough analytical methods and approaches based on software tools and simula
 tors.\n\nAI-Assisted Design-Space Analysis of High-Performance Arm Process
 ors\n\nThis work quantifies the impact of microarchitectural features in m
 odern high-performance Arm CPUs. To combat a parameter space that is too l
 arge to traverse naively, we employ a decision tree regression machine lea
 rning model to predict the number of execution cycles with 93.38% accuracy
  compared t...\n\n\nJoseph Moore, Tom Deakin, and Simon McIntosh-Smith (Un
 iversity of Bristol, England)\n---------------------\nUnderstanding VASP P
 ower Profiles on NVIDIA A100 GPUs\n\nPower is a critical limiting factor i
 n supercomputing as systems scale to exascale levels. To advance scientifi
 c computing, supercomputers must operate efficiently under limited power b
 udgets. Power-aware scheduling can help by enforcing power management stra
 tegies, but this requires a deep understa...\n\n\nZhengji Zhao, Brian Aust
 in, Ermal Rrapaj, and Nicholas Wright (Lawrence Berkeley National Laborato
 ry (LBNL))\n---------------------\nPerformance Analysis of Runtime Handlin
 g of Zero-Copy for OpenMP Programs on MI300A APUs\n\nIn current discrete G
 PU systems, the penalty of data movement between host and device memory is
  inevitable, forcing many large-scale applications to include optimization
 s that amortize this cost. On systems like the AMD Instinct™ MI300A series
  accelerators, based on the accelerated processing ...\n\n\nCarlo Bertolli
  (AMD, Inc.) and Thorsten Blass, Jan-Patrick Lehr, Doru Bercea, Dhruva Cha
 krabarti, Lynd Stringer, Nicole Aschenbrenner, Lawrence Meadows, and Ron L
 ieberman (AMD Research)\n---------------------\nComprehensive Performance 
 Modeling and System Design Insights for Foundation Models\n\nGenerative AI
 , in particular large transformer models, are increasingly driving HPC sys
 tem design in science and industry. We analyze performance characteristics
  of such transformer models and discuss their sensitivity to the transform
 er type, parallelization strategy, and HPC system features (accel...\n\n\n
 Shashank Subramanian, Ermal Rrapaj, and Peter Harrington (Lawrence Berkele
 y National Laboratory (LBNL)); Smeet Chheda (Stony Brook University); and 
 Steven Farrell, Brian Austin, Samuel Williams, Nicholas Wright, and Wahid 
 Bhimji (Lawrence Berkeley National Laboratory (LBNL))\n-------------------
 --\nAssessing the GPU Offload Threshold of GEMM and GEMV Kernels on Modern
  Heterogeneous HPC Systems\n\nWith an ever-growing compute advantage over 
 CPUs, GPUs are often used in workloads with ample BLAS computation to impr
 ove performance. However, several factors, including data-to-compute ratio
 , amount of data re-use, and data structure, can all impact performance. H
 ence, using a GPU is not a guarant...\n\n\nFinn Wilkinson, Alex Cockrean, 
 Wei-Chen Lin, Simon McIntosh-Smith, and Tom Deakin (University of Bristol,
  England)\n---------------------\nLLM-Inference-Bench: Inference Benchmark
 ing of Large Language Models on AI Accelerators\n\nLarge language models (
 LLMs) have propelled groundbreaking advancements across several domains an
 d are commonly used for text generation applications. However, the computa
 tional demands of these complex models pose significant challenges, requir
 ing efficient hardware acceleration. Benchmarking the p...\n\n\nKrishna Te
 ja Chitty-Venkata, Siddhisanket Raskar, Bharat Kale, Farah Ferdaus, Aditya
  Tanikanti, Ken Raffenetti, Valerie Taylor, Murali Emani, and Venkatram Vi
 shwanath (Argonne National Laboratory (ANL))\n---------------------\nPMBS2
 4 — Morning Break\n---------------------\nImpact of Varying BLAS Precision
  on DCMESH\n\nThe limiting factor in the application of high-accuracy quan
 tum molecular simulations to large systems has been the associated high co
 mputational costs in terms of both compute power and memory. In this paper
  we explore the use of various BLAS precision modes (BF16, TF32, and Compl
 ex_3M) in DCMESH, ...\n\n\nNariman Piroozan and John Pennycook (Intel Corp
 oration), Taufeq Razakh (University of Southern California (USC)), Peter C
 aday and Nalini Kumar (Intel Corporation), and Aiichiro Nakano (University
  of Southern California (USC))\n---------------------\nMicroarchitectural 
 Comparison and In-Core Modeling of State-of-the-Art CPUs: Grace, Sapphire 
 Rapids, and Genoa\n\nWith NVIDIA’s release of the Grace Superchip, all thr
 ee big semiconductor companies in HPC (AMD, Intel, NVIDIA) are currently c
 ompeting in the race for the best CPU. In this work we analyze the perform
 ance of these state-of-the-art CPUs and create an accurate in-core perform
 ance model for thei...\n\n\nJan Laukemann (University of Erlangen-Nurember
 g, Germany; Erlangen National High Performance Computing Center) and Georg
  Hager and Gerhard Wellein (Erlangen National High Performance Computing C
 enter)\n---------------------\nPonte Vecchio Across the Atlantic: Single-N
 ode Benchmarking of Two Intel GPU Systems\n\nIntel Data Center GPU Max 155
 0, known as Ponte Vecchio (PVC), is a new Intel GPU architecture for high-
 performance computing. It is the basis of two systems on the June 2024 TOP
 500 list, Dawn (#51) and Aurora (#2).\n\nThis work provides micro-benchmar
 king data on PVCs from which application developers...\n\n\nThomas Applenc
 ourt (Argonne National Laboratory (ANL)); Aditya Sadawarte (University of 
 Bristol, England); Servesh Muralidharan, Colleen Bertoni, JaeHyuk Kwack, Y
 e Luo, Esteban Rangel, John Tramm, and Yasaman Ghadar (Argonne National La
 boratory (ANL)); Arjen Tamerus and Chris Edsall (University of Cambridge, 
 England); and Tom Deakin (University of Bristol, England)\n---------------
 ------\nHello SME! Generating Fast Matrix Multiplication Kernels Using the
  Scalable Matrix Extension\n\nModern central processing units (CPUs) featu
 re single-instruction, multiple-data pipelines to accelerate compute-inten
 sive floating-point and fixed-point workloads. Traditionally, these pipeli
 nes and corresponding instruction set architectures (ISAs) were designed f
 or vector parallelism. The Scalabl...\n\n\nStefan Remke and Alexander Breu
 er (Friedrich Schiller University Jena, Germany)\n---------------------\nS
 ystem-Wide Roofline Profiling: A Case Study on NERSC’s Perlmutter Supercom
 puter\n\nHPC system architects routinely use application profiling and per
 formance modeling to evaluate hardware and software performance trade-offs
 . However, the focus on individual applications leaves gaps in the underst
 anding of system utilization because it is impractical to collect profiles
  and models f...\n\n\nBrian Austin, Dhruva Kulkarni, Brandon Cook, Samuel 
 Williams, and Nicholas Wright (Lawrence Berkeley National Laboratory (LBNL
 ))\n---------------------\nBenchmarking the Evolution of Performance and E
 nergy Efficiency Across Recent Generations of Intel Xeon Processors\n\nIn 
 the span of 1.5 years, Intel has launched four families of Xeon processors
 , with some novel architectural features. First, the Sapphire Rapids gener
 ation, which featured a version with on-package HBM; next, the Emerald Rap
 ids generation; and then Intel differentiated by releasing the performance
 -...\n\n\nIstvan Z. Reguly and Balázs Drávai (Pázmány Péter Catholic Unive
 rsity, Hungary)\n---------------------\nWorkload-Adaptive Scheduling for E
 fficient Use of Parallel File Systems in High-Performance Computing Cluste
 rs\n\nWhereas contentions within storage systems noticeably impact runtime
 s, shared bandwidth-type resources, such as Lustre, pose challenges for hi
 gh-performance computing cluster schedulers. Additionally, accurately esti
 mating job resource requirements, particularly related to I/O operations, 
 remains a ...\n\n\nAlexander Goponenko (University of Central Florida), Be
 njamin Allan and James Brandt (Sandia National Laboratories), and Damian D
 echev (University of Central Florida)\n---------------------\nPMBS24 — Aft
 ernoon Break\n---------------------\nPMBS24 — Lunch Break\n\nTag: Accelera
 tors, Modeling and Simulation, Performance Evaluation and/or Optimization 
 Tools\n\nRegistration Category: Workshop Reg Pass\n\nSession Chairs: Simon
  Hammond (National Nuclear Security Administration (NNSA)); Sascha Hunold 
 (Technical University of Vienna); and Steven A. Wright (University of York
 , England)
END:VEVENT
END:VCALENDAR
