BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143035Z
LOCATION:B302-B305
DTSTART;TZID=America/New_York:20241119T120000
DTEND;TZID=America/New_York:20241119T170000
UID:submissions.supercomputing.org_SC24_sess487@linklings.com
SUMMARY:Research, ACM SRC, and Doctoral Showcase Posters
DESCRIPTION:Evolving a Multi-Population Evolutionary-QAOA on Distributed Q
 PUs\n\nOur research combines an Evolutionary Algorithm with a Quantum Appr
 oximate Optimization Algorithm (QAOA) to update the ansatz parameters, in 
 place of traditional gradient-based methods, and benchmarks on the Max-Cut
  problem. We demonstrate that our Evolutionary-QAOA pairing performs on pa
 r or better...\n\n\nFrancesca Schiavello, Edoardo Altamura, and Stefano Me
 nsa (Hartree Centre, Science and Technology Facilities Council (STFC), UK)
 \n---------------------\nUncover the Overhead and Resource Usage for Handl
 ing KV Cache Overflow in LLM Inference\n\nLLM inference includes two phase
 s: prefill phase and decode phase. Prefill phase processes input tokens si
 multaneously to generate the first token. Decode phases generate the subse
 quent tokens one after another until they either meet a termination or rea
 ch the max length. To avoid recomputation, the...\n\n\nJie Ye (Illinois In
 stitute of Technology), Bogdan Nicolae (Argonne National Laboratory (ANL))
 , and Anthony Kougkas and Xian-He Sun (Illinois Institute of Technology)\n
 ---------------------\nEfficient Large Dynamic Graph Analysis on Emerging 
 Storage Technology\n\nGraph-structured data analysis is extensively used i
 n various real-world applications, including biology, social media, and re
 commendation systems. With the increasing prevalence of real-time data, ma
 ny graphs become dynamic and evolve over time. Thus, dynamic graph process
 ing systems become a neces...\n\n\nAbdullah Al Raqibul Islam (University o
 f North Carolina at Charlotte)\n---------------------\nEffects of Lossy Co
 mpression Data on Machine Learning Models\n\nMachine learning is a fundame
 ntal tool that is incorporated in fields across academia and industry. Due
  to the large amounts of data needed for training machine learning models,
  compression is utilized because it reduces the data footprint playing a c
 ritical role in storage. Machine learning involve...\n\n\nMax Faykus (Clem
 son University)\n---------------------\nI/O Characterization of Heterogene
 ous Workflows\n\nWorkflows consist of individual applications such as scie
 ntific simulations and data analytics. These applications constitute diffe
 rent stages of the workflow, each comprising heterogeneous characteristics
  such as run-times and system requirements. The heterogeneity in these wor
 kflow stages dictates...\n\n\nOlga Kogiou (Florida State University), Hari
 haran Devarajan and Chen Wang (Lawrence Livermore National Laboratory (LLN
 L)), Weikuan Yu (Florida State University), and Kathryn Mohror (Lawrence L
 ivermore National Laboratory (LLNL))\n---------------------\nFault-Toleran
 t Numerical Iterative Algorithms at Scale\n\nNumerical iterative algorithm
 s are struck by multiple error types when deployed on large-scale HPC plat
 forms: fail-stop errors (failures) and silent errors, striking both as com
 putation errors and memory bit-flips. Our novel approach provides efficien
 t fault-tolerant algorithms that are capable of d...\n\n\nAlix Tremodeux (
 ENS Lyon)\n---------------------\nBenchmarking Quantum-Inspired Optimizati
 on Platforms and Tools on an HPC Cluster\n\nThis work compares performance
  of various QUBO software tools on different platforms available on the So
 l supercomputer at Arizona State University. CPU, GPU and the NEC vector e
 ngine (VE) card provide various means to implement these computations on s
 imulated qubits while employing various solvers,...\n\n\nJencey Dole, Gil 
 Speyer, and Torey Battelle (Arizona State University)\n-------------------
 --\nNetCDFaster: A Geospatial Cyberinfrastructure Enhancing Multi-Dimensio
 nal Scientific Dataset Access and Visualization Through Machine Learning O
 ptimization\n\nThis project introduces an enhanced solution for accessing 
 and processing NetCDF data, a widely used standard in geosciences for stor
 ing multidimensional data. Existing tools often compromise on performance 
 or lack full workflow support. The proposed system integrates machine lear
 ning, specifically ...\n\n\nZhenlei Song and Zhe Zhang (Texas A&M Universi
 ty) and Alan Sussman (University of Maryland)\n---------------------\nScal
 able Performance and Accuracy Analysis for Distributed and Extreme-Scale S
 ystems\n\nThe Scalable Performance and Accuracy analysis for Distributed a
 nd Extreme-scale systems (SPADE) project focuses on advancing monitoring, 
 optimization, evaluation, and decision-making capabilities for extreme-sca
 le systems. This poster presents efforts targeting several advanced monito
 ring capabilit...\n\n\nHeike Jagode (University of Tennessee, Innovative C
 omputing Laboratory (ICL); University of Tennessee); Shirley Moore (Univer
 sity of Texas at El Paso); Vincent Weaver (University of Maine); Anthony D
 analis (University of Tennessee, Innovative Computing Laboratory (ICL)); a
 nd Christoph Lauter (University of Texas at El Paso)\n--------------------
 -\nInteractive and Tool-Agnostic ML-Driven Workflow for Automated HPC Perf
 ormance Modeling\n\nThis work presents an automated, reproducible, ML-base
 d performance modeling workflow for HPC systems. The proposed workflow ful
 ly automates data generation, preprocessing, ML model training and validat
 ion. Since the proposed approach is generic and not tailored to a specific
  application, our workfl...\n\n\nZhaobin Zhu and Sarah Neuwirth (Johannes 
 Gutenberg University Mainz)\n---------------------\nNeural Network Optimiz
 ation and Performance Analysis for Real-Time Object Detection at the Edge\
 n\nReal-time object detection is an important and computationally intensiv
 e task that is gaining more attention in the field of autonomous systems. 
 Recently, a novel object detection algorithm called RT-DETR has emerged, d
 emonstrating superior speed compared to the popular YOLO series. In recent
  years,...\n\n\nAnimesh Ghose (Brookhaven National Laboratory, Carnegie Me
 llon University)\n---------------------\nAnalyzing Alltoall Algorithms wit
 h SST\n\nAlltoall collective operations in MPI are critical in several typ
 es of computation, including matrix multiplication and transposition, and 
 machine learning applications. As a result, it is critical that these oper
 ations are performant for large amounts of data. Meanwhile, dragonfly netw
 orks are beco...\n\n\nShannon Kinkead (Sandia National Laboratories) and A
 manda Bienz (University of New Mexico)\n---------------------\nScalable Pl
 anning Platform for Orchestration of Autonomous Systems Across Edge-Cloud 
 Continuum\n\nEdge accelerators, such as NVIDIA Jetson, are enabling rapid 
 inference of deep neural network (DNN) models and computer vision algorith
 ms through low-end graphics processing unit (GPU) modules integrated with 
 ARM-based processors. Their compact form factor allows integration with mo
 bile platforms, s...\n\n\nSuman Raj (Indian Institute of Science, Bangalor
 e)\n---------------------\nLagrangian Particle-Tracking in GPU-Enabled Ext
 reme Scale Turbulence Simulations\n\nMany practical turbulent flow phenome
 na are naturally studied using a Lagrangian approach that treats the fluid
  medium as a collection of infinitesimal fluid particles. We present a GPU
 -accelerated algorithm for tracking particles in direct numerical simulati
 ons of isotropic turbulence, scaling up t...\n\n\nRohini Uma-Vaideswaran a
 nd P.K. Yeung (Georgia Institute of Technology) and Kiran Ravikumar (Scien
 ce and Technology Corporation, NASA Ames Research Center)\n---------------
 ------\nHPC Fastpass: Visualizing Descriptive and Predictive HPC Queue Tim
 e Data\n\nWith large HPC systems, users will often jockey for better queue
  times to get quicker results. Unfortunately, getting accurate estimations
  of queue times requires understanding complex and abundant data collected
  from myriad HPC system loggers. To aid with this, researchers are explori
 ng machine lea...\n\n\nConnor Scully-Allison (University of Utah, Scientif
 ic Computing and Imaging Institute (SCI); National Renewable Energy Labora
 tory (NREL))\n---------------------\nSWARM: Scientific Workflow Applicatio
 ns on Resilient Metasystem\n\nCurrent (centralized) resource management st
 rategies typically require a global view of distributed HPC systems, relyi
 ng on a cluster-wide resource manager for scheduling, with static, expert-
 tuned rules. This centralized decision-making approach suffers from resili
 ence, efficiency and scalability i...\n\n\nPawel Zuk (University of Southe
 rn California, Information Sciences Institute); Hongwei Jin (Argonne Natio
 nal Laboratory (ANL)); Imtiaz Mahmud (Lawrence Berkeley National Laborator
 y (LBNL)); Krishnan Raghavan (Argonne National Laboratory (ANL)); Komal Th
 areja (Renaissance Computing Institute (RENCI)); Shixun Wu (University of 
 California, Riverside); Prasanna Balaprakash (Oak Ridge National Laborator
 y (ORNL)); Franck Cappello (Argonne National Laboratory (ANL)); Zizhong Ch
 en (University of California, Riverside); Ewa Deelman (University of South
 ern California, Information Sciences Institute); Sheng Di (Argonne Nationa
 l Laboratory (ANL)); Aiden Hamade and Mariam Kiran (Oak Ridge National Lab
 oratory (ORNL)); Anirban Mandal, Erik Scott, and Cong Wang (Renaissance Co
 mputing Institute (RENCI)); and John Wu (Lawrence Berkeley National Labora
 tory (LBNL))\n---------------------\nA Zero-Copy Storage with Metadata-Dri
 ven File Management Using Persistent Memory\n\nPersistent Memory (PM) is a
  promising next-generation storage device, combining features of both vola
 tile memory (like DRAM) and non-volatile memory (like SSDs). Many studies 
 use PM to optimize training to advance deep learning technology. However, 
 these studies have not addressed the issue of multi...\n\n\nGuangxing Hu (
 North Carolina State University), Awais Khan (Oak Ridge National Laborator
 y (ORNL)), and Frank Mueller (North Carolina State University)\n----------
 -----------\nOptimal Client Selection Algorithms for Federated Learning\n\
 nDue to the heterogeneity of resources and data, client selection plays a 
 paramount role in the efficacy of Federated Learning (FL) systems. The tim
 e taken by a training round is determined by the slowest client. Also, ene
 rgy consumption and carbon footprint are seen as primary concerns. In this
  cont...\n\n\nAlan Nunes (Fluminense Federal University, Brazil; Universit
 y of Bordeaux, France); Cristina Boeres and Lúcia Drummond (Fluminense Fed
 eral University, Brazil); and Laércio Pilla (University of Bordeaux, Frenc
 h Institute for Research in Computer Science and Automation (INRIA))\n----
 -----------------\nMatRIS: Performance Portable Math Library of IRIS Runti
 me for Multi-Device Heterogeneity\n\nWe present the recent efforts for Mat
 RIS, the performance portable math library of IRIS runtime for multi-devic
 e heterogeneity. MatRIS provides dense linear algebra — BLAS (Basic Linear
  Algebra Subprograms) and LAPACK (Linear Algebra PACKage) — capabilities a
 cross different back ends ava...\n\n\nMohammad Alaul Haque Monil, Narasing
 a Rao Miniskar, Pedro Valero-Lara, Keita Teranishi, and Jeffrey S. Vetter 
 (Oak Ridge National Laboratory (ORNL))\n---------------------\nToward Perf
 ormance & Portability & Productivity in Parallel Programming\n\nAchieving 
 *performance*, *portability*, and *productivity* for data-parallel computa
 tions (e.g., MatMul and convolutions) has emerged as a major research chal
 lenge. The complex hardware design of contemporary parallel architectures,
  including GPUs and CPUs, requires advanced program optimizations to...\n\
 n\nAri Rasch (University of Münster, Germany)\n---------------------\nSees
 aw: Elastic Scaling for Task-Based Distributed Programs\n\nModern batch sc
 hedulers in HPC environments enable the shared use of available computatio
 nal resources via provisioning discrete sets of resources matching user re
 quirements. The lack of elasticity in such scenarios is often addressed us
 ing a Pilot job model where multiple separate requests are pool...\n\n\nMa
 tthew Chung (University of California, Riverside)\n---------------------\n
 Exploring DAOS as a Burst Buffer for a 100 Gbps DAQ Real-Time Streaming Sy
 stem\n\nWe present an experimental evaluation of a burst buffer for a real
 -time DAQ streaming system designed to transmit instrument data to remote 
 data centers. The system is based on EJ-FAT, a load balancing system capab
 le of Nx 100Gbps streams, distributing data from event sources to processi
 ng nodes. We...\n\n\nXinxin Mei, Graham Heyes, Vardan Gyurjyan, Carl Timme
 r, Michael Goodrich, Amitoj Singh, and Ilya Baldin (Thomas Jefferson Natio
 nal Accelerator Facility)\n---------------------\nMind Your Manners: Detox
 ifying Language Models via Attention Head Intervention\n\nTransformer-base
 d Large Language Models have advanced natural language processing with the
 ir ability to generate fluent text. However, these models exhibit and ampl
 ify toxicity and bias learned from training data, posing new ethical chall
 enges. We build upon the \attnlens{} framework to allow for sc...\n\n\nJor
 dan Pettyjohn (Colorado School of Mines)\n---------------------\nStalls an
 d Memory Analysis on Fujitsu A64FX and NVIDIA Grace\n\nARM-based multicore
  CPUs, such as NVIDIA Grace and Fujitsu A64FX, dominate contemporary HPC, 
 featuring 32-256 cores with cache hierarchies and up to 1 TB/s memory band
 width. While benchmarks like STREAM show similar performance across these 
 systems, diverse applications, particularly graph and neare...\n\n\nYan Ka
 ng (Pennsylvania State University), Sayan Ghosh (Pacific Northwest Nationa
 l Laboratory (PNNL)), Mahmut Kandemir (Pennsylvania State University), and
  Andres Marquez (Pacific Northwest National Laboratory (PNNL))\n----------
 -----------\nProfiling the Impact of Hyper-Threading on Pagosa Hydrocodes\
 n\nPagosa is a hydrodynamics computer code designed for massively parallel
  environments, operating within an Eulerian framework using a fixed Cartes
 ian mesh. This project investigates the performance of Pagosa on Sapphire 
 Rapids nodes, which feature many-core architectures and high bandwidth mem
 ory. By...\n\n\nBlake Milstead (University of Tennessee, Knoxville)\n-----
 ----------------\nEfficient, Scalable, Robust Neuromorphic High Performanc
 e Computing\n\nThe rapid advancement in Artificial Neural Networks (ANNs) 
 has paved the way for Spiking Neural Networks (SNNs), which offer signific
 ant advantages in energy efficiency and computational speed, especially on
  neuromorphic hardware. My research focuses on the development of Efficien
 t, Robust, and Scal...\n\n\nBiswadeep Chakraborty (Georgia Institute of Te
 chnology)\n---------------------\nPerformance of N10 Benchmarks with Diffe
 rent BLAS Implementations\n\nThe NERSC-10 Benchmark Suite is a collection 
 of tests which are designed to evaluate various aspects of system architec
 ture and performance. This study utilizes the suite to evaluate the CPU co
 mpute nodes of two leadership-class supercomputers, Perlmutter at NERSC an
 d Summit at OLCF, using two N10 b...\n\n\nShubh Pachchigar (NERSC at LBNL 
 (Lawrence Berkeley National Laboratory), San Francisco State University) a
 nd Adam Lavely, Rahulkumar Gayatri, and Brandon Cook (NERSC at LBNL (Lawre
 nce Berkeley National Laboratory))\n---------------------\nA Novel Gradien
 t Compression Design with Ultra-High Compression Ratio for Communication-E
 fficient Federated Learning\n\nFederated learning is a privacy-preserving 
 machine learning approach. It allows numerous geographically distributed c
 lients to collaboratively train a large model while maintaining local data
  privacy. In heterogeneous device settings, limited network bandwidth is a
  major bottleneck that constrains s...\n\n\nZhijing Ye and Xiaodong Yu (St
 evens Institute of Technology)\n---------------------\nEmpowering Scientif
 ic Datasets with Large Language Models\n\nThe growing volume and complexit
 y of scientific data pose significant challenges in data management, organ
 ization, and analysis. Our objective is to enhance the utilization of hist
 orical scientific datasets across various disciplines. To address this, we
  propose integrating large language models (LL...\n\n\nBrian Chen and Alvi
 n Hoang (University of California, Riverside; Pacific Northwest National L
 aboratory (PNNL))\n---------------------\nJACC: HPC Meta-Programming and P
 erformance Portability Ecosystem for Julia Language\n\nWe present JACC (Ju
 lia for ACCelerators), the first high-level, meta-programming, and perform
 ance-portable model for the just-in-time and LLVM-based Julia language. JA
 CC provides a unified and lightweight front end across different back ends
  available in Julia, enabling the same Julia code to run ef...\n\n\nPedro 
 Valero-Lara, William Godoy, Het Mankad, Keita Teranishi, and Jeffrey Vette
 r (Oak Ridge National Laboratory (ORNL))\n---------------------\nEnhancing
  the Traditional Benchmarks for Parallel Computing Education\n\nThis work 
 introduces enhanced benchmark suites — NeoRodinia, RedOptBench and CUDAMic
 roBench — specifically designed to enrich the educational landscape of par
 allel programming. By integrating practical examples and detailed optimiza
 tion processes into traditional benchmarks, these suites...\n\n\nXinyao Yi
  (University of North Carolina at Charlotte), Anjia Wang (Intel Corporatio
 n), and Yonghong Yan (University of North Carolina at Charlotte)\n--------
 -------------\nCoVA: Compiler for Versatile Architectures\n\nThe rapid adv
 ancement in computing demands and the increasing complexity of modern appl
 ications (e.g., image processing, numerical computation, and machine learn
 ing) necessitate efficient heterogeneous computing solutions. Project CoVA
  proposes an MLIR-based compilation flow designed to bridge the g...\n\n\n
 Hongzheng Tian (University of California, Irvine; Hewlett Packard Labs); A
 lok Mishra (Hewlett Packard Labs); and Sitao Huang (University of Californ
 ia, Irvine)\n---------------------\nComputational Radiation Hydrodynamics 
 with FleCSI\n\nPhotons are essential in many matter-light interaction syst
 ems, with applications in physics (e.g., core-collapse supernovae) and eng
 ineering (e.g., neutron transport). Although the theory of radiating fluid
 s has been established since the 1980s, developing robust and efficient nu
 merical simulations...\n\n\nMammadbaghir Baghirzade (The University of Tex
 as at Austin; CCS-7, Los Alamos National Laboratory); Md Shihab Shahriar K
 han (Michigan State University; CCS-7, Los Alamos National Laboratory); Yo
 onsoo Kim (California Institute of Technology; CCS-7, Los Alamos National 
 Laboratory); Sudarshan Neopane (University of Tennessee, Knoxville; CCS-7,
  Los Alamos National Laboratory); Alexander Strack (University of Stuttgar
 t, Germany; CCS-7, Los Alamos National Laboratory); Farhana Taiyebah (Flor
 ida State University; CCS-7, Los Alamos National Laboratory); and Julien L
 oiseau and Hyun Lim (Los Alamos National Laboratory (LANL))\n-------------
 --------\nDisCostiC: Simulating MPI Applications Without Executing Code\n\
 nWe present the cross-architecture parallel simulation framework DisCostiC
  (Distributed Cost in Clusters). It predicts the performance of real or hy
 pothetical, massively parallel MPI(+X) programs on current and future supe
 rcomputer systems. The novelty of DisCostiC is that it employs analytical,
  firs...\n\n\nAyesha Afzal, Georg Hager, and Gerhard Wellein (Friedrich-Al
 exander University, Erlangen-Nuremberg; Erlangen National High Performance
  Computing Center)\n---------------------\nIntegrating HPCToolkit with Too
 ls for Automated Analysis\n\nHPCToolkit enables users to gather detailed i
 nformation about application performance. Users can capture fine-grained m
 easurement data, which may include instruction-level samples on CPUs and G
 PUs. Collected data can be huge, making manual inspection using GUI tools 
 difficult and time-consuming. We ...\n\n\nDragana Grbic (Lawrence Livermor
 e National Laboratory (LLNL), Rice University); Matthew LeGendre and Olga 
 Pearce (Lawrence Livermore National Laboratory (LLNL)); and John Mellor-Cr
 ummey (Rice University)\n---------------------\nGoing Beyond the Chicken a
 nd Egg Situation with Modern MPI Features\n\nMPI has become the de facto s
 tandard for distributed memory computing since its inception in 1994. Whil
 e the MPI standard has evolved to include new technologies like RDMA, many
  applications still rely on the original set of MPI operations.\n\nThis th
 esis investigates the current usage of MPI. We note...\n\n\nTim Jammer (Te
 chnical University Darmstadt)\n---------------------\nLM-Offload: Performa
 nce Model-Guided Generative Inference of Large Language Models with Parall
 elism Control\n\nLarge language models (LLMs) have achieved remarkable suc
 cess in various natural language processing tasks. However, LLM inference 
 is highly computational and memory-intensive, creating extreme deployment 
 challenges. Tensor offloading, combined with tensor quantization and async
 hronous task executio...\n\n\nJianbo Wu (University of California, Merced)
 ; Jie Ren (College of William & Mary); Shuangyan Yang (University of Calif
 ornia, Merced); Konstantinos Parasyris, Giorgis Georgakoudis, and Ignacio 
 Laguna (Lawrence Livermore National Laboratory (LLNL)); and Dong Li (Unive
 rsity of California, Merced)\n---------------------\nLarge-Scale Randomize
 d Program Generation with Large Language Models​\n\nLarge, diverse dataset
 s of executable programs are required for training and running machine lea
 rning models to find insights in program performance. While many open-sour
 ce code repositories exist freely on popular software development websites
  such as GitHub, the safety and executability of such pr...\n\n\nTyler Lam
  (University of California, Berkeley) and Lingda Li (Brookhaven National L
 aboratory)\n---------------------\nQ-NFSO: Exploring Quantum Applications,
  Noise Management, Fault Injection, Resource Scheduling and Optimization i
 n the NISQ Era\n\nQuantum computing has achieved significant milestones in
  recent years, underscoring its potential benefits for NP-hard application
 s both currently and in the future. Despite these advancements, contempora
 ry quantum computers are hindered by noise and a limited number of qubits.
  These limitations pos...\n\n\nBetis Baheri (Kent State University)\n-----
 ----------------\nSanQus: Staleness and Quantization-Aware Full-Graph Dece
 ntralized Training in GNNs\n\nGraph neural networks (GNNs) have demonstrat
 ed significant success in modeling graphs; however, they encounter challen
 ges in efficiently scaling to large graphs. To address this, we propose th
 e SanQus system, advancing our previous work, Sancus. SanQus reduces the n
 eed for expensive communication am...\n\n\nHongbo Yin (Hong Kong Universit
 y of Science and Technology, Guangzhou); Jingshu Peng (Hong Kong Universit
 y of Science and Technology); Qiyu Liu (Southwest University); Zhao Chen (
 Hong Kong University of Science and Technology); Yingxia Shao (Beijing Uni
 versity of Posts and Telecommunications); Yanyan Shen (Shanghai Jiao Tong 
 University); Lei Chen (Hong Kong University of Science and Technology); an
 d Jiannong Cao (The Hong Kong Polytechnic University)\n-------------------
 --\nPoseidon: A Source-to-Source Translator for Holistic HPC Optimization 
 of Ocean Models on Regular Grids\n\nOcean simulation models often underper
 form on modern high-performance computing (HPC) architectures, necessitati
 ng costly and time-consuming code rewrites.\n\nWe introduce Poseidon, an H
 PC-oriented source-to-source translator for Fortran-based fluid dynamics s
 olvers used in ocean and weather models wi...\n\n\nMaurice Brémond (French
  Institute for Research in Computer Science and Automation (INRIA); Univer
 sité Grenoble Alpes, France); Hugo Brunie (Université Grenoble Alpes, Fran
 ce; French Institute for Research in Computer Science and Automation (INRI
 A)); Laurent Debreu (French Institute for Research in Computer Science and
  Automation (INRIA); Université Grenoble Alpes, France); Rupert W. Ford (H
 artree Centre, Science and Technology Facilities Council (STFC), UK); Flor
 ian Lemarié (French Institute for Research in Computer Science and Automat
 ion (INRIA); Université Grenoble Alpes, France); Anna Mittermair (Technica
 l University of Munich); Andrew R. Porter (Hartree Centre, Science and Tec
 hnology Facilities Council (STFC), UK); Julien Rémy (French Institute for 
 Research in Computer Science and Automation (INRIA); Université Grenoble A
 lpes, France); Philippe Rosales and Martin Schreiber (Université Grenoble 
 Alpes, France; French Institute for Research in Computer Science and Autom
 ation (INRIA)); Martin Schulz (Technical University Munich); Sergi Siso (H
 artree Centre, Science and Technology Facilities Council (STFC), UK); and 
 Arthur Vidard (French Institute for Research in Computer Science and Autom
 ation (INRIA); Université Grenoble Alpes, France)\n---------------------\n
 Active Learning for Metamaterial Optimization on HPC and QC Integrated Sys
 tems\n\nActive learning algorithms, integrating machine learning, quantum 
 computing and optics simulation in an iterative loop, offer a promising ap
 proach to optimizing metamaterials. However, these algorithms can face dif
 ficulties in optimizing highly complex structures due to computational lim
 itations. Hi...\n\n\nSeongmin Kim and In-Saeng Suh (Oak Ridge National Lab
 oratory (ORNL))\n---------------------\nAn Accurate and Scalable Multidime
 nsional Quantum Solver for Partial Differential Equations\n\nQuantum compu
 ting is an innovative technology that can solve certain problems faster th
 an classical computing. One of its promising applications is in solving pa
 rtial differential equations (PDEs). However, current PDE solvers that are
  based on variational-quantum-eigensolver (VQE) techniques suffer...\n\n\n
 Manu Chaudhary, Ishraq Islam, Alvir Nobel, Dylan Kneidel, Vinayak Jha, Ben
  Phillips, Kareem El-Araby, Manish Singh, and Esam El-Araby (University of
  Kansas)\n---------------------\nImproving SpGEMM Performance Through Reor
 dering and Cluster-Wise Computation\n\nSparse Matrix-Matrix Multiplication
  (SpGEMM) is a key kernel in many scientific applications and graph worklo
 ads. SpGEMM is known to suffer from poor performance due to irregular memo
 ry access patterns. Gustavson's algorithm, a traditional approach for SpGE
 MM, involves row/column-wise operations, fa...\n\n\nAbdullah Al Raqibul Is
 lam (University of North Carolina at Charlotte), Helen Xu (Georgia Institu
 te of Technology), Dong Dai (University of Delaware), and Aydin Buluç (Law
 rence Berkeley National Laboratory (LBNL))\n---------------------\nBreakin
 g the Barriers to Effective Supercomputing: Web Dashboard for Job Accounti
 ng and Performance Metrics\n\nThe NSF-funded Anvil supercomputer, built an
 d maintained by the Purdue University Rosen Center for Advanced Computing,
  enables efficient research computing across a variety of scientific domai
 ns nationwide. In addition to the traditional terminal interface, Anvil us
 es the open-source web portal fram...\n\n\nRichie Tan (Purdue University),
  Anjali Rajesh (Rutgers University), and Guangzhen Jin and Rajesh Kalyanam
  (Purdue University)\n---------------------\nBringing It HOME: Analyzing C
 ontention Hotspots Across the Memory Hierarchy with Low Overhead\n\nThe in
 creasing demand for computing in scientific research has given rise to mem
 ory contention and performance bottlenecks. Existing solutions often carry
  high overheads or lack the necessary detail for effective contention miti
 gation. To tackle these challenges, we are developing a powerful tool, H..
 .\n\n\nDhruv Gajaria and Andrés Márquez (Pacific Northwest National Labora
 tory (PNNL))\n---------------------\nSupporting End Users in Implementing 
 Quantum Computing Applications\n\nQuantum computing applications require e
 xpert knowledge to perform complex steps: (1) selecting a suitable quantum
  algorithm, (2) generating the quantum circuit, (3) compiling/executing th
 e quantum circuit, and (4) decoding the results. This creates high entry b
 arriers for end users with limited exp...\n\n\nNils Quetschlich (Technical
  University of Munich)\n---------------------\nMachine Learning Applicatio
 ns for Early-Stage Ovarian Cancer Diagnosis\n\nOvarian cancer (OC) signifi
 cantly impacts women's health, and despite its prevalence, remains without
  a definitive cure. Early detection is crucial for improving treatment out
 comes and reducing mortality rates and healthcare system costs. Leveraging
  advancements in machine learning, our study seeks ...\n\n\nDelina Mekonne
 n and Kazem Cheshmi (McMaster University, Ontario, Canada)\n--------------
 -------\nSimplifying HPC Resource Selection: A Tool for Optimizing Executi
 on Time and Cost on Azure\n\nAzure Cloud offers a wide range of resources 
 for running HPC workloads, requiring users to configure their deployment b
 y selecting VM types, number of VMs, and processes per VM. Suboptimal deci
 sions may lead to longer execution times or additional costs for the user.
  We are developing an open-source...\n\n\nMarco A. S. Netto, Wolfgang De S
 alvador, and Davide Vanzo (Microsoft Corporation)\n---------------------\n
 QFw: A Quantum Framework for Large-Scale HPC Ecosystems\n\nThis work exten
 ds Quantum Framework (QFw) by integrating it with Northwest Quantum Simula
 tor (NWQ-Sim) and by introducing a lightweight python library (qfw\_backen
 d) that allows multiple frontends (e.g., Qiskit) to interact with QFw. Thi
 s extension enables QFw to flexibly decouple frontends from bac...\n\n\nSr
 ikar Chundury (North Carolina State University); Amir Shehata, Thomas Naug
 hton, and Seongmin Kim (Oak Ridge National Laboratory (ORNL)); Frank Muell
 er (North Carolina State University); and In-Saeng Suh (Oak Ridge National
  Laboratory (ORNL))\n---------------------\nHigh-Performance Computing Res
 ilience Analysis Using Large Language Models\n\nThis doctoral showcase hig
 hlights three pivotal works conducted during my PhD that collectively adva
 nce the field of high-performance computing (HPC) resilience analysis usin
 g large language models (LLMs).\n\nThe first work introduces HAPPA, a modu
 lar platform for HPC Application Resilience Analysis. ...\n\n\nHailong Jia
 ng (Kent State University)\n---------------------\nEnhancing HPC I/O Perfo
 rmance: Leveraging Runtime and Offline I/O Optimization Frameworks\n\nThe 
 existing parallel I/O stack is complex and difficult to tune due to the in
 terdependencies among multiple factors that impact the performance of data
  movement between storage and compute systems. When performance is slower 
 than expected, end-users, developers, and system administrators rely on I/
 ...\n\n\nHammad Bin Ather (University of Oregon)\n---------------------\nG
 PU Compression (for Scientific Data) Done Right\n\nError-bounded lossy com
 pression is a critical technique for significantly reducing scientific dat
 a volumes. Compared to CPU-based compressors, GPU-based compressors exhibi
 t substantially higher throughputs, fitting better for today's HPC applica
 tions. To overcome the data challenge, GPU-based scient...\n\n\nJiannan Ti
 an (Indiana University), Xin Liang (University of Kentucky), and Sheng Di 
 and Franck Cappello (Argonne National Laboratory (ANL))\n-----------------
 ----\n5G in Practice: Measuring Emerging Wireless Technology in Rural Iowa
  for Edge Devices in Distributed Computation Workloads\n\nEdge computing, 
 the notion of moving computational tasks from servers to the data-generati
 ng network edge, is an increasingly popular model for data processing. 5G 
 wireless technologies offer an opportunity to enable complex distributed e
 dge computing workflows by minimizing the overhead incurred in...\n\n\nZac
 k Murry and Alicia Esquivel Morel (University of Missouri, Columbia; Unive
 rsity of Chicago) and Kate Keahey (Argonne National Laboratory (ANL), Univ
 ersity of Chicago)\n---------------------\nPrompt Phrase Ordering Using La
 rge Language Models in HPC: Evaluating Prompt Sensitivity\n\nLarge languag
 e models (LLMs) often require well-designed prompts for effective response
 s, but optimizing prompts is challenging due to prompt sensitivity, where 
 small changes can cause significant performance variations. This study eva
 luates prompt performance across all permutations of independent ...\n\n\n
 Noah Thomasson (Oak Ridge National Laboratory (ORNL), PCIP Internship Prog
 ram) and Hilda Klasky (Oak Ridge National Laboratory (ORNL))\n------------
 ---------\nEnhancing HPC Resource Management to Integrate Quantum Workflow
 s\n\nIn recent years, quantum computing has demonstrated the potential to 
 revolutionize specific algorithms and applications by solving problems exp
 onentially faster than classical computers. However, its widespread adopti
 on for general computing remains a future prospect. In this work we demons
 trate the...\n\n\nAmir Shehata, Thomas Naughton, and In-Saeng Suh (Oak Rid
 ge National Laboratory (ORNL))\n---------------------\nCluster Management 
 with Containerization on Switches\n\nThis research explores the utilizatio
 n of underused computational resources in network switches from Arista and
  Mellanox by deploying containers for auxiliary tasks in computational clu
 ster networks. We tested five scenarios: 1) running cloud-init services fo
 r efficient boot processes, 2) using Tele...\n\n\nRobin Lee Simpson (Los A
 lamos National Laboratory (LANL); California Polytechnic State University,
  San Luis Obispo); Anvitha Ramachandran (Los Alamos National Laboratory (L
 ANL), University of Southern California (USC)); and Dohyun Lee (Los Alamos
  National Laboratory (LANL), University of Illinois Urbana-Champaign)\n---
 ------------------\nA Survey-Based Evaluation of the Efficacy of a Girls W
 ho Code Club at the University of Southern Indiana\n\nPrograms like Girls 
 Who Code (GWC) are pivotal in working to inspire and equip young women wit
 h the skills and confidence needed to pursue careers in computing. Underst
 anding the impact of such initiatives is particularly important for addres
 sing the decline in interest among girls aged 13 to 17, a ...\n\n\nMaya Se
 shan, Cathy Sandoval, Alyson Collins, and Srishti Srivastava (University o
 f Southern Indiana)\n---------------------\nPrototype Development and Test
 ing of a Smart Buoy System for Coastal and Marine Ecosystems Using IBIS\n\
 nEffective environmental monitoring traditionally required physical presen
 ce at research sites. IoT technologies now enable remote data collection, 
 revolutionizing this field. This work presents a prototype of smart buoys 
 designed for coastal and marine ecosystems. Located in the IBIS testbed, t
 hese ...\n\n\nJulia Harper (Loyola University, Chicago)\n-----------------
 ----\nFault Tolerance in Krylov Subspace Methods\n\nToday’s complex HPC sy
 stems are incredibly powerful yet equally likely to experience failures. T
 he scientific applications on these HPC systems are mostly iterative in na
 ture. Iterative solvers have some inherent fault tolerance, but they are s
 till susceptible to errors. One subset of these it...\n\n\nSandesh Pandit 
 (University of Alabama in Huntsville)\n---------------------\nJUmPER: Perf
 ormance Data Monitoring, Instrumentation and Visualization for Jupyter Not
 ebooks\n\nComputational performance, e.g. CPU or GPU utilization, is cruci
 al for analyzing machine learning (ML) applications and their resource-eff
 icient deployment. However, the ML community often lacks accessible tools 
 for holistic performance engineering, especially during exploratory progra
 mming such as ...\n\n\nElias Werner and Anton Rygin (Center for Scalable D
 ata Analytics and Artificial Intelligence (ScaDS.AI) Dresden/Leipzig; Tech
 nical University Dresden, Center for Interdisciplinary Digital Sciences (C
 IDS)); Andreas Gocht-Zech (German Environment Agency, AI Lab); and Sebasti
 an Döbel and Matthias Lieber (Technical University Dresden, ZIH; Technical
  University Dresden, Center for Interdisciplinary Digital Sciences (CIDS))
 \n---------------------\nHARVEST-2.0: High-Performance Vision Framework fo
 r End-to-End Preprocessing, Training, Inference, and Visualization\n\nDeep
  learning (DL) thrives on data; however, it inherits a major limitation: t
 raining and testing datasets must be fully annotated for supervised deep n
 eural networks (DNNs) training. To address this challenge, we introduce HA
 RVEST-2.0, a high-performance computer-vision framework for end-to-end dat
 ...\n\n\nNawras Alnaasan, Anirudh Potlapally, Tian Chen, Matthew Lieber, A
 amir Shafi, Hari Subramoni, Scott Shearer, and Dhabaleswar K. Panda (The O
 hio State University)\n---------------------\nPower Patterns: Understandin
 g the Energy Dynamics of I/O for Parallel Storage Configurations\n\nAs HPC
  applications become more I/O intensive, understanding their power consump
 tion patterns is necessary to develop energy-saving solutions. Here, we ev
 aluate the energy consumption of I/O operations on two popular HPC paralle
 l file systems: Lustre and DAOS. We develop models to predict the energy..
 .\n\n\nMaya Purohit (University of Chicago, Colby College) and Valerie Hay
 ot-Sasson and Kyle Chard (University of Chicago)\n---------------------\nO
 n the Accuracy and Efficiency of Approximate Triangle Counting via Randomi
 zed Numerical Linear Algebra\n\nWe study two algorithmic approaches to app
 roximate triangle counting and compare their accuracy and efficiency. The 
 first one is based on randomized matrix-matrix multiplication, which can b
 e faster, simpler, and more parallelizable on modern processors. The secon
 d is based on trace estimation, whic...\n\n\nIrene Simó Muñoz, Koby Hayash
 i, and Richard Vuduc (Georgia Institute of Technology)\n------------------
 ---\nComparing Cache Utilization Trends for Regional Scientific Caches wit
 h Transfer Learning Models\n\nTo enhance data sharing and reduce access la
 tency in scientific collaborations, high energy physics (LHC CMS experimen
 t) employs regional in-network storage caches. Accurate predictions of cac
 he utilization trends help design new caching policies and improve capacit
 y planning. This study leverages t...\n\n\nErica Wang (California Institut
 e of Technology, Lawrence Berkeley National Laboratory (LBNL))\n----------
 -----------\nScalable Low-Latency Hardware Function Chaining with Chain Co
 ntrol Circuit\n\nTo deliver advanced services with high performance and fl
 exibility, we are developing a computing platform that integrates various 
 hardware accelerators (HWAs). Our goal is to build customized services for
  each user by combining application functions. To achieve this, we propose
  Hardware Function Ch...\n\n\nYuta Ukon, Takahiro Kawahara, Yuki Arikawa, 
 Naoki Miura, and Teruaki Ishizaki (NTT Corporation); Wataru Kanemori, Ryok
 o Tamura, and Kenjiro Mori (Fujitsu Ltd); and Takeshi Sakamoto (NTT Corpor
 ation)\n---------------------\nAn Adaptive Kernel Execution for Dynamic Ap
 plications on GPUs Using CUDA Graphs\n\nWe propose a novel approach for ex
 ecuting dynamic applications on GPUs. Different from traditional approache
 s that use a single kernel, our method allows the GPU to autonomously allo
 cate computational resources at runtime. We decompose a kernel into multip
 le fragment kernels and dynamically launch a...\n\n\nKento Kitamura, Kenji
  Tanaka, and Kazunori Seno (NTT Corporation)\n---------------------\nPerfo
 rmance of Inline Compression with Software Caching for Reducing the Memory
  Footprint in pySDC\n\nThe volume of data required for high performance co
 mputing (HPC) jobs is growing faster than the memory storage available to 
 store the required data, leading to performance bottlenecks. Hence the nee
 d for inline data compression, which reduces the amount of allocated memor
 y needed by storing all dat...\n\n\nEmily Lattanzio (Clemson University)\n
 ---------------------\nAccelerating Communications in High-Performance Sci
 entific Workflows\n\nAdvances in networks, accelerators, and cloud service
 s encourage programmers to reconsider where to compute — such as when fast
  networks make it cost-effective to compute on remote accelerators despite
  added latency. Workflow and cloud-hosted serverless computing frameworks 
 can manage multi-st...\n\n\nJ. Gregory Pauloski (University of Chicago, Ar
 gonne National Laboratory (ANL))\n---------------------\nFormal Approaches
  to Characterize Emerging Arithmetic Realizations\n\nAs HPC models grow in
 creasingly complex, disparities in floating point implementations across h
 ardware platforms begin to pose significant challenges to reproducibility 
 and reliability. This is especially so, given that HPC employs hardware op
 timized for performance, which quite often deviates from ...\n\n\nPaul Jia
 ng (Purdue University) and Iana Lin (University of California, Berkeley)\n
 ---------------------\nAccelerating HPC Workflow Results and Performance R
 eproducibility Analytics\n\nModern high-performance computing (HPC) workfl
 ows produce massive datasets, often exceeding 100+ TB per day, driven by i
 nstruments collecting data at gigabytes per second. These workflows, execu
 ted on advanced HPC systems with heterogeneous storage devices, high-perfo
 rmance microprocessors, accelera...\n\n\nKevin Assogba (Rochester Institut
 e of Technology)\n---------------------\nScored Non-Deterministic Finite A
 utomata Processor for Sequence Alignment\n\nWith the surge in symbolic dat
 a across various fields, efficient pattern matching and regular expression
  processing have become crucial. Non-deterministic Finite Automata (NFA), 
 commonly used for pattern matching, face memory bottlenecks on general-pur
 pose platforms. This has driven interest in Doma...\n\n\nRyan Karbowniczak
  and Rasha Karakchi (University of South Carolina)\n---------------------\
 nDART-X: Software Infrastructure for Prototyping In-Memory Data Transfer B
 etween Ensemble Data Assimilation and Coupled Earth Systems Models\n\nData
  assimilation (DA) integrates real-world data into coupled climate models,
  enhancing prediction accuracy and capturing Earth system complexity. NCAR
 ’s Data Assimilation Research Testbed (DART) is an ensemble DA tool for cl
 imate predictions with NCAR’s Community Earth System Model (CE...\n\n\nAnh
  Pham (Mount Holyoke College, National Center for Atmospheric Research (NC
 AR)); Suman Shekhar (Rutgers University, National Center for Atmospheric R
 esearch (NCAR)); and Helen Kershaw, Dan Amrhein, and Ufuk Turuncoglu (Nati
 onal Center for Atmospheric Research (NCAR))\n---------------------\nData 
 Layout Optimizations for Tensor Applications\n\nThe performance of tensor 
 applications is often bottlenecked by data movement across the memory subs
 ystem. This dissertation contributes domain-specific programming framework
 s (compilers and runtime systems) that optimize data movement in tensor ap
 plications. We develop novel execution reordering an...\n\n\nMahesh Lakshm
 inarasimhan (University of Utah)\n---------------------\nKVSort: Drastical
 ly Improving LLM Inference Performance via KV Cache Compression\n\nLarge l
 anguage model (LLM) deployment necessitates high inference throughput due 
 to the increasing demand for text generation. To accelerate inference, the
  prefill mechanism avoids repeated computations via introducing KV Cache (
 in HBM). However, the KV cache size increases with the input and genera...
 \n\n\nBaixi Sun (Indiana University); Dingwen Tao (Institute of Computing 
 Technology, Chinese Academy of Sciences); Xiaodong Yu (Stevens Institute o
 f Technology); and Fengguang Song (Indiana University)\n------------------
 ---\nExploring Software-Defined Networking for Routing in Dragonfly Topolo
 gy\n\nWhile efficient routing in Dragonfly networks presents significant c
 hallenges, the advent of Software-defined Networking (SDN) offers new oppo
 rtunities for routing optimization by providing a global network view. Thi
 s research proposes an SDN-based adaptive routing approach for Dragonfly i
 nterconnec...\n\n\nRam Sharan Chaulagain and Xin Yuan (Florida State Unive
 rsity)\n---------------------\nGeneralizing ExaDigiT Datacenter Digital Tw
 in Framework for Multiple Architectures\n\nExaDigiT is an open-source fram
 ework for developing comprehensive digital twins (DTs) of liquid-cooled su
 percomputers. DTs merge telemetry, simulations, and AI/ML/RL to create a c
 omplete virtual representation of a system. This provides an effective too
 l for testing a variety of system optimizations...\n\n\nSedrick Bouknight 
 and Matthias Maiterth (Oak Ridge National Laboratory (ORNL)); Tim Dykes (H
 ewlett Packard Enterprise (HPE)); Rashadul Kabir (Colorado State Universit
 y); and Vineet Kumar, Scott Greenwood, Jesse Hines, Wes Brewer, and Feiyi 
 Wang (Oak Ridge National Laboratory (ORNL))\n---------------------\nThe P3
  Explorer: Exploring the Performance, Portability, and Productivity Wilder
 ness\n\nThis paper documents the development of a web-based tool designed 
 to organise and present visual representations of performance, portability
 , and productivity (P3) data from previously published scientific studies.
  The P3 Explorer operates as both an open repository of scientific data an
 d a data das...\n\n\nMatthew A. Smith and Steven A. Wright (University of 
 York, England) and Zaman Lantra and Gihan Mudalige (University of Warwick,
  England)\n---------------------\nAssessing Matrix Multiplication Performa
 nce with Fully Homomorphic Encryption\n\nAs data from various domains is i
 ncreasingly shared and processed in the cloud, Homomorphic Encryption (HE)
  provides a crucial solution for ensuring privacy in the post-quantum era.
  In this work, we evaluate the performance and accuracy of the HE matrix m
 ultiplication leveraging SEAL\nlibrary kernels...\n\n\nJusto Molina, Rocío
  Carratalá-Sáez, Manuel F. Dolz, and Sandra Catalán (Jaume I University, S
 pain)\n---------------------\nTrackable Agent-Based Evolution Models at Wa
 fer Scale\n\nEmerging ML/AI hardware accelerators, like the 850,000 proces
 sor Cerebras Wafer-Scale Engine (WSE), hold great promise to scale up the 
 capabilities of evolutionary computation. However, challenges remain in ma
 intaining visibility into underlying evolutionary processes while efficien
 tly utilizing the...\n\n\nMatthew Andres Moreno and Connor Yang (Universit
 y of Michigan), Emily Dolson (Michigan State University), and Luis Zaman (
 University of Michigan)\n---------------------\nParallel Verification of N
 eural Networks Applied to Medical Imaging\n\nNeural network verification p
 rovides model robustness guarantees in the presence of noise. We generate 
 verification specifications for medical imaging models based on the U-Net 
 architecture and solve pixel-by-pixel verification problems on a massive s
 cale. Efficiency of this NP-complete problem is s...\n\n\nJonathan Andreas
 en (Georgia Tech Research Institute), Diego Manzanas Lopez and Taylor John
 son (Vanderbilt University), and Yatis Dodia (Georgia Tech Research Instit
 ute)\n---------------------\nAssessing the Impact of Real-Time Traffic Upd
 ates on Traffic Flow: A High-Performance Computing Perspective on Scalabil
 ity and Demand\n\nIn the dynamic landscape of smart cities, it is necessar
 y to improve the potential of real-time traffic data and high-performance 
 computing to optimize traffic flow through dynamic re-routing strategies. 
 Our research contributes to the assessment of how real-time traffic optimi
 zation and alternative...\n\n\nPaulo Silva, Pavlína Smolková, Sofia Michai
 lidu, Jakub Beránek, Roman Macháček, Kateřina Slaninová, and Jan Martinovi
 č (IT4Innovations, VSB - Technical University of Ostrava) and Radim Cmar (
 Sygic a.s.)\n---------------------\niSeeMore: Design of a 256-Node RPi Clu
 ster to Visualize LLM Computation Through Light and Movement for Mass Audi
 ences\n\niSeeMore is a kinetic cluster of 256-Raspberry Pi (RPi) computers
  that visually realizes supercomputing concepts (parallelism, data flow, a
 lgorithms) through servo and LED driven movement. In this poster, we descr
 ibe the design of iSeeMore, the first large-scale cluster to combine movem
 ent and light...\n\n\nHayden Estes, Gregory Bolet, Collin Dunphy, Matt Lor
 enzo, Mallika Pamula, Kalina Kazmierczak, Sahithi Sarva, Margaret Ellis, S
 am Blanchard, Godmar Back, and Kirk Cameron (Virginia Tech)\n-------------
 --------\nGNN-RL: An Intelligent HPC Resource Scheduler\n\nEfficient resou
 rce allocation in high-performance computing (HPC) environments is crucial
  for optimizing utilization, minimizing make-span, and enhancing throughpu
 t. We propose GNN-RL, a novel intelligent scheduler that leverages a hybri
 d Graph Neural Network and Reinforcement Learning model, learni...\n\n\nKy
 rian Chinemeze Adimora and Hongyang Sun (University of Kansas)\n----------
 -----------\nCommunication Hiding for Matrix-Free Finite Element Operators
  of a Complex PDE: Nonlinear Stokes Flow of Earth’s Mantle\n\nFor large-sc
 ale matrix-free finite-element PDE solvers, parallel matrix-vector product
 s typically comprise the dominant computational cost. Synchronization step
 s for input and output, when degrees of freedom (DOFs) lying on inter-proc
 ess boundaries are communicated, become the dominant serial portio...\n\n\
 nMax Heldman and Johann Rudi (Virginia Tech)\n---------------------\nProfi
 ling Communication Overhead in 3D Parallel Pretrain of Large Language Mode
 ls\n\nTraining large language models (LLMs) efficiently requires addressin
 g the communication overhead introduced by parallelism strategies like Ten
 sor, Pipeline, and Data Parallelism. This work profiles the communication 
 patterns in LLM pretraining using the Polaris supercomputer, highlighting 
 the impact...\n\n\nWeijin Liu and Xiaodong Yu (Stevens Institute of Techno
 logy)\n---------------------\nParallelization of the Finite Element-Based 
 Mesh Warping Algorithm Using Hybrid Parallel Programming\n\nWarping large 
 volume meshes is computationally expensive and has applications to biomech
 anics, aerodynamics, image processing, and cardiology. Existing parallel i
 mplementations of mesh warping algorithms do not take advantage of shared-
 memory and one-sided communication features available in MPI-3. ...\n\n\nA
 bir Haque (University of Kansas)\n---------------------\nEdge-Enabled Real
 -Time Data Processing in Power-Efficient Weather Stations Using IBIS\n\nTh
 ere is a growing need to acquire a larger quantity of meteorological data 
 to address climate change. In this paper, we design an improved Automatic 
 Weather Station (AWS) based on a prototype from the National Center for At
 mospheric Research (NCAR). We integrate this weather station with IBIS, a 
 pl...\n\n\nWilliam Fowler (Tufts University); Alicia Esquivel Morel (Unive
 rsity of Missouri, Columbia); and Kate Keahey (University of Chicago, Argo
 nne National Laboratory (ANL))\n---------------------\nORCHA: A Performanc
 e Portability System for Flash-X — A Multiphysics Application Software\n\n
 Heterogeneous platforms demand tailored data structures, algorithm mapping
 s, and efficient execution, often leading to numerous, hard-to-maintain va
 riants of source codes. ORCHA, our performance portability orchestration s
 ystem, addresses these challenges through abstractions and code generation
 , st...\n\n\nYoungjun Lee and Wesley Kwiecinski (Argonne National Laborato
 ry (ANL)); Klaus Weide (Argonne National Laboratory (ANL), University of C
 hicago); Johann Rudi (Virginia Tech); Jared O’Neal (Argonne National Labor
 atory (ANL)); Emil Vatai and Mohamed Wahib (RIKEN); and Tom Klosterman and
  Anshu Dubey (Argonne National Laboratory (ANL))\n---------------------\nE
 fficient Approaches to Analyzing Large Dynamic Networks\n\nDynamic graphs,
  characterized by their evolving topologies over time, necessitate continu
 ous updates to their associated graph properties, including shortest paths
 , vertex coloring, and strongly connected components. Traditional static g
 raph algorithms, which re-compute properties following each mod...\n\n\nAr
 indam Khanda and S.M. Shovan (Missouri University of Science and Technolog
 y), Aashish Pandey and Farahnaz Hosseini (University of North Texas), Saja
 l K. Das (Missouri University of Science and Technology), Boyana Norris (U
 niversity of Oregon), and Sanjukta Bhowmick (University of North Texas)\n-
 --------------------\nExploiting Data Compression and Low Precision for Ex
 ascale Fusion Turbulence Simulations\n\nFuture exascale systems will featu
 re unprecedented computing power with over 10^18 FLOPS, provided by thousa
 nds of heterogeneous computing nodes. To fully harness such potential, app
 lications must scale effectively and efficiently utilize these computing a
 rchitectures.\n\nMany performance losses on la...\n\n\nJan Laukemann (Univ
 ersity of Erlangen-Nuremberg, Germany; Erlangen National High Performance 
 Computing Center); Felix Jung (Technical University of Munich); Carl-Marti
 n Pfeiler (Max Planck Institute for Plasma Physics); Diego Jimenez (Max Pl
 anck Computing and Data Facility); Carsten Clauss (ParTec AG, Germany); Ti
 lman Dannert (Max Planck Computing and Data Facility); Frank Jenko (Max Pl
 anck Institute for Plasma Physics); Erwin Laure (Max Planck Computing and 
 Data Facility); Martin Schulz (Technical University of Munich); and Gerhar
 d Wellein (Erlangen National High Performance Computing Center)\n---------
 ------------\nAlgorithmic and Optimization Techniques for Graph Applicatio
 ns in Heterogeneous Systems at Scale\n\nAs heterogeneity becomes commonpla
 ce in HPC systems, algorithmic and optimization techniques are needed to a
 ddress the challenges that come with it, especially for irregular applicat
 ions. This includes workload balancing, scheduling, latency tolerance, and
  memory utilization and contention, among ot...\n\n\nReece Neff (North Car
 olina State University, Pacific Northwest National Laboratory (PNNL))\n---
 ------------------\nMemory Disaggregation in Serverless Computing\n\nIn re
 cent years, the slowing of advancements in memory technology and applicati
 ons’ increasing demand for memory have resulted in high performance comput
 ation becoming bottlenecked by availability of memory. One existing soluti
 on, far memory, involves swapping pages to a remote machine rather ...\n\n
 \nJoshua Lowe (Illinois Institute of Technology, Rose-Hulman Institute of 
 Technology) and Nanda Velugoti (Illinois Institute of Technology)\n-------
 --------------\nCreating Code LLMs for HPC: It’s LLMs All the Way Down\n\n
 Large language models (LLMs) are being used increasingly by software devel
 opers, researchers, and students to assist them in coding tasks. While new
 er LLMs have been improving their coding abilities with regards to serial 
 coding tasks, they consistently perform worse when it comes to parallelism
  and...\n\n\nAman Chaturvedi, Daniel Nichols, Siddharth Singh, and Abhinav
  Bhatele (University of Maryland)\n---------------------\nImprovement of B
 ridges-2 Resource Utilization Through User Optimization\n\nThis poster pre
 sents our two-phase solution for improving GPU utilization in NSF-funded A
 CCESS high-performance computing (HPC) clusters, with a pilot implementati
 on on Pittsburgh Supercomputing Center’s Bridges-2. Our approach addresses
  the limitations of Open XdMoD, which lacks per-job GPU u...\n\n\nWalter A
 shworth (University of Tennessee, Knoxville); Julian Uran (Pittsburgh Supe
 rcomputing Center (PSC), Carnegie Mellon University); Michela Taufer (Univ
 ersity of Tennessee, Knoxville); and Paola Buitrago (Pittsburgh Supercompu
 ting Center (PSC), Carnegie Mellon University)\n---------------------\nPer
 sistent and Partitioned MPI for Stencil Communication\n\nMany parallel app
 lications rely on iterative stencil operations, whose performance is domin
 ated by communication costs at large scales. Several MPI optimizations, su
 ch as persistent and partitioned communication, reduce overheads and impro
 ve communication efficiency through amortized setup costs and...\n\n\nGera
 ld Collom and Amanda Bienz (University of New Mexico)\n-------------------
 --\nPerformance of LAMMPS-SNAP in Different Runtime Environments\n\nLAMMPS
  (Large-scale Atomic/Molecular Massively Parallel Simulator) is a widely u
 sed molecular dynamics simulator. It is used here to simulate the high-pre
 ssure BC8 phase of carbon using the Spectral Neighbor Analysis Potential (
 SNAP). This simulation employs the Kokkos C++ performance portability la..
 .\n\n\nShubh Pachchigar (San Francisco State University, NERSC at LBNL (La
 wrence Berkeley National Laboratory)) and Adam Lavely, Rahulkumar Gayatri,
  and Brandon Cook (NERSC at LBNL (Lawrence Berkeley National Laboratory))\
 n---------------------\nPINE: Efficient Yet Effective Piecewise Linear Tre
 es\n\nDecision trees are popularly used in statistics and machine learning
 . Piecewise linear trees, a type of model-based decision tree, employ line
 ar models to evaluate splits and predict outcomes at the leaf nodes. While
  they can offer high accuracy, they are computationally expensive, and cur
 rently, no...\n\n\nZecheng Li and Jiajia Li (North Carolina State Universi
 ty)\n---------------------\nDesigning Efficient Data Reduction Approaches 
 for Multi-Resolution Simulations on HPC Systems\n\nAs supercomputers advan
 ce towards exascale capabilities, computational intensity increases signif
 icantly, and the volume of data requiring storage and transmission experie
 nces exponential growth. Multi-resolution methods, such as Adaptive Mesh R
 efinement (AMR), have emerged as an effective solution ...\n\n\nDaoce Wang
  (Indiana University)\n---------------------\nDesign of Reliable and Effic
 ient Syscall Hooking Library for a Parallel File System\n\nTo provide file
  systems in user space, FUSE and syscall interception libraries have been 
 used. However, FUSE has a performance degradation problem, and the syscall
  interception library has problems with reliability and portability. \n\nT
 his study proposes a design and implementation of the reliable an...\n\n\n
 Haruka Miyauchi (University of Tsukuba, Japan)\n---------------------\nPro
 filing and Bottleneck Identification for Large Language Model Optimization
 s\n\nLarge language models (LLMs) have shown they can perform scientific t
 asks. They are capable of assisting researchers in data interpretation, in
 strument operation, knowledge synthesis, and hypothesis generation. Howeve
 r, LLMs must first be trained on a large dataset of scientific tasks and d
 ata. Trai...\n\n\nAlvin Hoang and Brian Chen (University of California, Ri
 verside; Pacific Northwest National Laboratory (PNNL))\n------------------
 ---\nEstablishing Best Practices for Applying Inline Compressed Arrays to 
 Improve Performance in HPC\n\nHPC applications require massive amounts of 
 memory to process large datasets. While data compression is used to avoid 
 bottlenecks in transmission and storage, it is still necessary to decompre
 ss this data into memory to use it. Inline compressed arrays (ICA) is a me
 thod which keeps the data compress...\n\n\nAdam Whitestone (Clemson Univer
 sity)\n---------------------\nIncreasing the Efficiency of Neutral Atoms b
 y Reducing Qubit Waste from Measurement-Related Ejections\n\nQuantum compu
 ting is an emerging field that has had an impact on various domains. This 
 poster focuses on the advantageous neutral atom technology for quantum com
 puting. In neutral atom technology, a significant challenge arises in the 
 measurement of quantum output, where the need to eject atoms phys...\n\n\n
 Jason Ludmir and Tirthak Patel (Rice University)\n---------------------\nA
  Sparse Approach for Translation-Based Training of Knowledge Graph Embeddi
 ngs\n\nKnowledge graph (KG) learning offers a powerful framework for gener
 ating new knowledge and making inferences. Training KG embedding can take 
 a significantly long time, especially for larger datasets. Our analysis sh
 ows that the gradient computation of embedding and vector normalization ar
 e the domin...\n\n\nMd Saidul Hoque Anik and Ariful Azad (Indiana Universi
 ty)\n---------------------\nPcMINER: Mining Performance-Related Commits at
  Scale\n\nPerformance inefficiencies in software can severely impact appli
 cation quality and resource utilization. Addressing these issues often req
 uires significant developer effort, yet the lack of large-scale, open-sour
 ce performance datasets hinders the development of effective mitigation st
 rategies. To f...\n\n\nMd Abul Kalam Azad, Manoj Alexender, Matthew Alexen
 der, Syed Salauddin Mohammad Tariq, Foyzul Hassan, and Probir Roy (Univers
 ity of Michigan - Dearborn)\n---------------------\nDevelopment of TEZip i
 n PyTorch: Integrating New Prediction Models into an Existing Compression 
 Framework\n\nIn high performance computing, researchers often work with ex
 tremely large time-series datasets. Compression techniques allow them to s
 hrink their data down, allowing for quicker and less storage-intensive tra
 nsfers. TEZip is a compression model that utilizes PredNet (a video predic
 tion model) to pr...\n\n\nAkshay Nambudiripad (Columbia University), Amarj
 it Singh and Kento Sato (RIKEN), and Weikuan Yu (Florida State University)
 \n---------------------\nTrusted Platform Provisioning for the OpenCHAMI C
 luster Management Stack\n\nHigh performance computing (HPC) clusters have 
 traditionally relied on proprietary provisioning and management infrastruc
 ture. This can be problematic, especially with regard to ongoing security 
 and maintenance for vendored systems.\n\nAs an alternative to this, the Lo
 s Alamos National Laboratory (LAN...\n\n\nLucas Ritzdorf (Los Alamos Natio
 nal Laboratory (LANL), University of Washington) and Nicholas Jones, David
  Allen, and Travis Cotton (Los Alamos National Laboratory (LANL))\n-------
 --------------\nTowards Scalable Quantum Simulation on Wafer-Scale Engines
 \n\nQuantum computing poses many benefits to the computing world, given it
 s efficiency over classical computers in problem spaces. However, the high
  noise and low logical qubit count of contemporary noisy-intermediate-scal
 e-quantum (NISQ) devices make the development and execution of quantum alg
 orithms ...\n\n\nBen Phillips, Dylan Kneidel, Alvir Nobel, and Esam El-Ara
 by (University of Kansas)\n---------------------\nImproving the Performanc
 e of Proof-of-Space in Blockchain Systems\n\nBlockchain technologies enabl
 e the success of digital currencies by providing security, decentralizatio
 n, and trustless operation. Two dominant consensus algorithms, Bitcoin’s P
 roof-of-Work and Ethereum’s Proof-of-Stake, balance security, scalability,
  and energy efficiency, though PoW is...\n\n\nVarvara Bondarenko (Illinois
  Institute of Technology), Renato Diaz (University of Central Florida), an
 d Lan Nguyen (Illinois Institute of Technology)\n---------------------\nFF
 T-Based Spherical Harmonics and Radial Transforms on GPU\n\nModern high-pe
 rformance computing clusters switch to the GPUs, as opposed to CPUs, as th
 e source of their computational power. GPUs are tailored for data-parallel
  algorithms where multiple cores perform the same operations on different 
 memory locations. However, making CPU code run within GPU constr...\n\n\nD
 mitrii Tolmachev (ETH Zürich)\n---------------------\nMeteorologic Real-Ti
 me Extreme Learning Machine for Pressure Prediction\n\nSignificant advance
 s in weather prediction stemmed primarily from combining observational dat
 a, sophisticated modeling techniques, and analysis of simulated or histori
 cal weather data. Additionally, the applicability of machine learning-base
 d applications on edge devices has expanded, addressing a v...\n\n\nAnagha
  Ram (University of Chicago); Kate Keahey (University of Chicago, Argonne 
 National Laboratory (ANL)); and Alicia Esquivel Morel (University of Misso
 uri, Columbia)\n---------------------\nWeb-Based Simulator of Superscalar 
 RISC-V Processors\n\nUnlock the power of superscalar processor design with
  our cutting-edge RISC-V simulator! Tailored for IT students, researchers,
  and HPC professionals, this web-based tool brings complex architectures t
 o life with an intuitive, customizable interface. Explore processor compon
 ents, tweak configuration...\n\n\nJiri Jaros (Brno University of Technolog
 y, Faculty of Information Technology) and Michal Majer, Jakub Horky, and J
 an Vavra (Brno University of Technology)\n---------------------\nAn Error-
 Bounded Lossy Compression Method with Bit-Adaptive Quantization for Partic
 le Data\n\nWe present error-bounded lossy compression tailored for particl
 e datasets from diverse scientific applications in cosmology, fluid dynami
 cs, and fusion energy sciences. As today's high-performance computing capa
 bilities advance, these datasets often reach trillions of points, posing s
 ignificant anal...\n\n\nCongrong Ren (The Ohio State University), Sheng Di
  (Argonne National Laboratory (ANL)), Longtao Zhang and Kai Zhao (Florida 
 State University), and Hanqi Guo (The Ohio State University)\n------------
 ---------\nEnhancing Performance Reproducibility on HPC Workflows\n\nRepro
 ducibility is related to achieving consistent performance across multiple 
 runs of the same application in an identical computing environment. From a
  computational perspective, jobs should run for equal time, with equal per
 formance repeatedly. Without this, we see significant variations in perfo.
 ..\n\n\nAinara Garcia (Clemson University, Argonne National Laboratory (AN
 L))\n---------------------\nPerfFlowAspect: A User-Friendly Performance To
 ol for Scientific Workflows\n\nIn scientific computing, artificial intelli
 gence and/or machine learning (AI/ML) are appearing increasingly often in 
 scientific workflows on supercomputers, due to their ability to solve more
  complex problems. With respect to performance analysis tools, the nature 
 of these workflows creates new requ...\n\n\nAliza Lisan (University of Ore
 gon), Tapasya Patki and Stephanie Brink (Lawrence Livermore National Labor
 atory (LLNL)), Zhiwei Yang and Spencer Greene (Worcester Polytechnic Insti
 tute (WPI)), Konstantinos Parasyris (Lawrence Livermore National Laborator
 y (LLNL)), Shubbhi Taneja (Worcester Polytechnic Institute (WPI)), and Han
 k Childs (University of Oregon)\n---------------------\nPipeInfer: Acceler
 ating LLM Inference Using Asynchronous Pipelined Speculation​\n\nInference
  of large language models (LLMs) across computer clusters has become a foc
 al point of research in recent times, with many acceleration techniques ta
 king inspiration from CPU speculative execution. These techniques reduce m
 emory bandwidth requirements, but also increase\nlatency per inference...\
 n\n\nBranden Butler and Sixing Yu (Iowa State University), Arya Mazaheri (
 Technical University Darmstadt), and Ali Jannesari (Iowa State University)
 \n---------------------\nLarge Genomic Language Models: Towards Their Hype
 rparameter Optimization\n\nThis study explores hyperparameter optimization
  for encoder-only genomics large language models, balancing machine learni
 ng performance, hardware resource usage and power consumption. Multiple se
 arch techniques and objective functions were implemented, and from the res
 ults we obtained two types of m...\n\n\nAndrei Politov and Neringa Jurenai
 te (Technical University Dresden, ZIH; Center for Scalable Data Analytics 
 and Artificial Intelligence (ScaDS.AI) Dresden/Leipzig)\n-----------------
 ----\nHardware-Independent Sampling Library for CPUs and (Multi-)GPUs: hws
 \n\nTo be energy efficient and fully utilize modern hardware, it is import
 ant to gain as much insight as possible into the performance and efficienc
 y of an application. Especially in the age of artificial intelligence, it 
 gets more and more important to keep, e.g., track of the total energy cons
 umption ...\n\n\nMarcel Breyer (University of Stuttgart, Germany; Institut
 e for Parallel and Distributed Systems) and Alexander Van Craen and Dirk P
 flüger (University of Stuttgart, Germany; University of Stuttgart, Institu
 te for Parallel and Distributed Systems)\n---------------------\nExplorati
 on of Super-Resolution Techniques for Image Compression\n\nTEZIP is a (de)
 compression framework leveraging PredNet, a deep neural network designed f
 or video prediction tasks, to exploit temporal locality in time-evolving d
 ata. This study evaluates video super-resolution (VSR) models, which enhan
 ce low-resolution images by reconstructing high-resolution ones...\n\n\nAl
 exis Amoyo (RIKEN, Florida State University); Amarjit Singh and Kento Sato
  (RIKEN); and Weikuan Yu (Florida State University)\n---------------------
 \nAlgorithmic Patterns from Computational Biology for Proxy Application De
 velopment and Co-Design\n\nHigh-performance computing hardware is co-devel
 oped with U.S. DOE codes, and proxy applications (apps) based on these cod
 es are critical technologies for iterative innovation. Numerical modeling/
 simulation proxies have been the most impactful in co-design. To broaden t
 he types of computation availab...\n\n\nLogan Williams and Gavin Conant (N
 orth Carolina State University) and Jan Ciesko and Amy Powell (Sandia Nati
 onal Laboratories)\n---------------------\nCluster-Based Methodology for C
 haracterizing the Performance of Portable Applications\n\nThis work focuse
 s on performance portability and proposes a methodological approach to ass
 essing and explaining how different kernels behave across various hardware
  architectures using the RAJA Performance Suite (RAJAPerf). Our methodolog
 y leverages metrics from the Intel top-down pipeline and clust...\n\n\nBef
 ikir Bogale (University of Tennessee, Knoxville); Olga Pearce and Tom Scog
 land (Lawrence Livermore National Laboratory (LLNL)); and Michela Taufer (
 University of Tennessee, Knoxville)\n---------------------\nCharacterizing
  the Performance of the GENE-X Code for Gyrokinetic Turbulence Simulations
 \n\nSimulating plasma turbulence in the edge region of a magnetic confinem
 ent fusion (MCF) device is crucial for identifying optimal operational sce
 narios for future fusion energy commercialization. GENE-X, an Eulerian ele
 ctromagnetic gyrokinetic code, can simulate plasma turbulence throughout a
 n MCF de...\n\n\nJordy Trilaksono (Max Planck Institute for Plasma Physics
 , Tokamak Theory Division); Jeremy Johnathan Williams (KTH Royal Institute
  of Technology); Philipp Ulbl (Max Planck Institute for Plasma Physics, To
 kamak Theory Division); Tilman Dannert and Erwin Laure (Max Planck Computi
 ng and Data Facility); Stefano Markidis (KTH Royal Institute of Technology
 ); and Frank Jenko (Max Planck Institute for Plasma Physics, Tokamak Theor
 y Division)\n---------------------\nTurbocharging Dask Apps: Accelerating 
 Data Flow with ProxyStore\n\nDespite advancements in distributed computing
  libraries, performance challenges, such as data serialization and transfe
 r, still persist. We focus on understanding data limitations within Dask, 
 a versatile and popular Python library designed for distributed and parall
 el computing, and then investigat...\n\n\nKlaudiusz Rydzy (Loyola Universi
 ty, Chicago)\n---------------------\nPerformance Engineering and Mesoscale
 -Microscale Coupling for Wind Energy Simulations\n\nWind farm simulations 
 require data from mesoscale atmospheric simulations as initial and boundar
 y conditions for the microscale turbine environments. The Energy Research 
 and Forecasting (ERF) code bridges this scale gap and provides an efficien
 t GPU-enabled parallel implementation with adaptive mesh...\n\n\nMukul Dav
 e, Ann Almgren, Donald Willcox, Weiqun Zhang, and Aaron Lattanzi (Lawrence
  Berkeley National Laboratory (LBNL))\n---------------------\nMIGnificient
 : Fast, Isolated, and GPU-Enabled Serverless Functions\n\nEmerging applica
 tions in machine learning and personalized medicine introduce new challeng
 es and requirements for secure computing. While the exclusive allocation o
 f resources to a single tenant provides necessary isolation, it also comes
  at the cost of hardware underutilization. While solutions lik...\n\n\nMar
 cin Copik and Alexandru Calotoiu (ETH Zürich); Pengy Zhou (University of T
 oronto, Canada); Lukas Tobler (AYES); and Torsten Hoefler (ETH Zürich)\n--
 -------------------\nQuantum Volume Benchmarking Simulators on HPC Systems
 \n\nQuantum computing simulators via classical computing are essential for
  small-scale comprehension and testing of quantum computing algorithms. Qu
 antum Volume (QV) is a well-established benchmark for comparing Quantum Pr
 ocessing Units (QPUs) in the NISQ era. However, there is no QV benchmark o
 f a larg...\n\n\nChristian Boehme (GWDG, Germany); Lourens van Niekerk (GW
 DG, Germany; Georg-August Universität Göttingen (GWDG)); Dhiraj Kumar and 
 Aasish Kumar Sharma (University of Göttingen, GWDG, Germany); and Tino Mei
 sel and Martin Leandro Paleico (GWDG, Germany)\n---------------------\nNew
  Semi-Implicit Electrostatic Particle-In-Cell Method to Extend Scope of th
 e Exascale WarpX Code\n\nTime-explicit fully-kinetic simulation of magneti
 cally confined fusion plasmas is out of reach of existing supercomputers d
 ue to the multi-scale nature of the system. Physics approximations and tim
 e-implicit methods are typically used to simulate such plasmas. A new ener
 gy-conserving, semi-implicit ...\n\n\nRoelof E. Groenewald and Daniel C. B
 arnes (TAE Technologies); Mikhail Tyushev (University of Saskatchewan, Can
 ada); Weiqun Zhang and Axel Huebl (Lawrence Berkeley National Laboratory (
 LBNL)); Ales Necas (TAE Technologies); Andrei Smolyakov (University of Sas
 katchewan, Canada); Jean-Luc Vay (Lawrence Berkeley National Laboratory (L
 BNL)); Toshiki Tajima (TAE Technologies; University of California, Irvine)
 ; and Sean Dettrick (TAE Technologies)\n---------------------\nPredicting 
 Dataset Popularity for Improved Distributed Content Caching in High Energy
  Physics\n\nIn High Energy Physics (HEP), large-scale experiments generate
  massive amounts of data that are distributed globally. To reduce redundan
 t data transfers and improve analysis efficiency, a disk caching system na
 med XCache is used to manage data accesses. By analyzing 11 months of acce
 ss logs (4.5 mil...\n\n\nMalavikha Sudarshan (University of California, Be
 rkeley)\n---------------------\nExploring Fine-Grained Memory Analysis for
  PIM Offloading\n\nIn modern computing, a challenge is the data bottleneck
  between the CPU and RAM. This issue arises because the CPU can process da
 ta faster than it can be accessed from RAM; this is worsened by the fact t
 hat large amounts of RAM are less accessible than a powerful CPU. Furtherm
 ore, RAM’s high c...\n\n\nThomas Papka (Loyola University, Chicago) and Na
 nda Velugoti and Kyle Hale (Illinois Institute of Technology)\n-----------
 ----------\nRAPIDS: Reduced API Data-Transfer Specifications\n\nPerformanc
 e of all-purpose communication libraries like MPI is fundamentally limited
  by the all-in-one approach these libraries use. RAPIDS (Reduced API Data-
 transfer Specifications) divides the functionality of these libraries into
  separate, more focused APIs, enabling library and application devel...\n\
 n\nRiley Shipley and Anthony Skjellum (Tennessee Technological University)
 , Purushotham Bangalore (University of Alabama), and Patrick Bridges (Univ
 ersity of New Mexico)\n---------------------\nGenerating Coupled Cluster C
 ode for Modern Distributed Memory Tensor Software\n\nScientific groups are
  struggling to adapt their codes to quickly-developing GPU-based HPC platf
 orms. The domain of distributed coupled cluster (CC) calculations is not a
 n exception. Moreover, our applications to tiny QED effects require higher
 -order CC which include thousands of tensor contractions,...\n\n\nJan Bran
 dejs, Johann Pototschnig, and Trond Saue (Laboratoire de Chimie et Physiqu
 e Quantiques)\n---------------------\nIdentifying Regions of Non-Determini
 sm in HPC Simulations Through Event Graph Alignment\n\nHigh performance co
 mputing (HPC) applications using MPI (Message Passing Interface) often fac
 e non-determinism (ND) due to asynchronous MPI calls, making ND source ide
 ntification challenging. Modeling execution as an event graph, where MPI c
 alls are nodes and communication is edges, can be useful. F...\n\n\nDhroov
  Pandey (University of North Texas)\n---------------------\nEve: Less Memo
 ry, Same Might\n\nAdaptive optimizers, which adjust the learning rate for 
 individual parameters, have become the standard for training deep neural n
 etworks. AdamW is a popular adaptive method that maintains two optimizer s
 tate values (momentum and variance) per parameter, doubling the model’s me
 mory usage durin...\n\n\nAditya Tomar (University of California, Berkeley)
  and Siddharth Singh, Tom Goldstein, and Abhinav Bhatele (University of Ma
 ryland)\n---------------------\nScalable Motif Counting on Large-Scale Dyn
 amic Graphs\n\nMotifs, small subgraphs of k vertices, such as triangles an
 d cliques, are well studied for static networks. They are used to characte
 rize different biological networks and align different networks. Counting 
 motifs reveal insights into the topological structure of a network such as
  MPI event graphs. ...\n\n\nAli Khan (University of North Texas)\n--------
 -------------\nFAS-GED: GPU-Accelerated Graph Edit Distance Computation\n\
 nGraph Edit Distance (GED) is a fundamental metric for assessing graph sim
 ilarity with critical applications across various domains, including bioin
 formatics, classification, and pattern recognition. However, the exponenti
 al computational complexity of GED has hindered its adoption for large-sca
 le gr...\n\n\nAdel Dabah and Andreas Herten (Forschungszentrum Jülich)\n--
 -------------------\nQDD: Multi-Node Implementation of Decision Diagram-Ba
 sed Quantum Circuit Simulator with Ring Communication and Auto SWAP Insert
 ion\n\nThis poster introduces QDD, a multi-node implementation of a decisi
 on diagram (DD)-based quantum circuit simulator. DD-based simulators offer
  faster simulation of algorithms like Shor's compared to statevector (SV) 
 simulators by compressing the quantum state using a graph representation. 
 However, pa...\n\n\nYusuke Kimura and Junpei Koyama (Fujitsu Ltd) and Shao
 wen Li, Hiroyuki Sato, and Masahiro Fujita (The University of Tokyo, Japan
 )\n---------------------\nFortranX: Harnessing Code Generation, Portabilit
 y, and Heterogeneity in Fortran\n\nDue to its historical popularity, Fortr
 an was used to implement many important scientific applications. The compl
 exity of these applications, along with the transition to modern high perf
 ormance languages like C++, has made modernization and optimization challe
 nging for these applications. Significa...\n\n\nSanil Rao (Carnegie Mellon
  University); Mike Franusich (SpiralGen, Inc.); Mohammad Alaul Haque Monil
 , Het Mankad, and Jeffrey Vetter (Oak Ridge National Laboratory (ORNL)); a
 nd Franz Franchetti (Carnegie Mellon University)\n---------------------\nB
 enchmarking and Modeling of Producer-Consumer Data Movement Performance in
  Scientific Workflows\n\nIn this poster, we present the Analytics4X (A4X) 
 framework, a workflow framework that enables systematic studies, evaluatio
 n, and modeling of HPC/HTC workflow data movement in a flexible and contro
 lled environment. We apply A4X to an in situ molecular dynamics workflow t
 o assess the performance trad...\n\n\nIan Lumsden (University of Tennessee
 ); Olga Pearce, Jae-Seung Yeom, and Tom Scogland (Lawrence Livermore Natio
 nal Laboratory (LLNL)); and Michela Taufer (University of Tennessee, Knoxv
 ille)\n---------------------\nAI-Based Scalable Analytics for Improving Pe
 rformance and Resilience of HPC Systems\n\nAs high-performance computing (
 HPC) advances to exascale levels, its role in scientific fields such as me
 dicine, climate research, finance, and scientific computing becomes increa
 singly critical. However, these large-scale systems are susceptible to per
 formance variations caused by anomalies, includ...\n\n\nEfe Sencan, Beste 
 Oztop, Ayse Coskun, Brian Kulis, and Manuel Egele (Boston University) and 
 Benjamin Schwaller, Vitus Leung, and Jim Brandt (Sandia National Laborator
 ies)\n---------------------\nImproving Polyhedral-Based Optimizations with
  Dynamic Coordinate Descent\n\nPolyhedral optimizations have been a corner
 stone of kernel optimization for many years. These techniques use a geomet
 ric model of loop iterations to enable transformations like tiling, fusion
 , and fission. The elegance of this approach lies in its ability to produc
 e highly efficient code through ful...\n\n\nGaurav Verma (Stony Brook Univ
 ersity; Stony Brook University, Institute for Advanced Computational Scien
 ce (IACS)); Michael Canesche (Universidade Federal de Minas Gerais, Brazil
 ); Barbara Chapman (Stony Brook University; Stony Brook University, Instit
 ute for Advanced Computational Science (IACS)); and Fernando Magno Quintão
  Pereira (Universidade Federal de Minas Gerais, Brazil)\n-----------------
 ----\nA Comparison Study of Open Source LLMs for HPC Ticket Answering​\n\n
 We are designing an automatic ticket answering service for computing cente
 rs such as Texas Advanced Computing Center (TACC), National Center for Sup
 ercomputing Applications (NCSA), and San Diego Supercomputer Center (SDSC)
 . In this work, we investigate the capability and feasibility of open sour
 ce l...\n\n\nFangru Linghu, Mingkai Zheng, and Zhao Zhang (Rutgers Univers
 ity)\n---------------------\nProposal for a Parallel Automatic Tuning Usin
 g d-Spline According to the Operating State of the Computer System\n\nSoft
 ware auto-tuning (AT) is a technology that parameterizes factors that affe
 ct performance as "performance parameters" and automatically tunes perform
 ance. In AT, the tool searches for better values for performance parameter
 s by repeatedly running the program. Therefore, if the target is a program
 ...\n\n\nYuga Yajima, Akihiro Fujii, and Teruo Tanaka (Kogakuin University
 , Japan)\n\nRegistration Category: Tech Program Reg Pass, Exhibits Reg Pas
 s\n\nSession Chairs: Ayesha Afzal (Friedrich-Alexander-Universität Erlange
 n-Nürnberg, Erlangen National High Performance Computing Center); Sally El
 lingson (University of Kentucky); and Alan Sussman (University of Maryland
 )
END:VEVENT
END:VCALENDAR
