All times are EDT (UTC/GMT -04 hours)
Speaker/Presenting Author in Italics
Session Chair/Session Host: J. Kepner & A. Reuther
- Keynote Talk: Expanding Quantum: MIT’s Quantum Initiative
- Danna Freedman (MIT)
Organizers: D. Tiwari & J. Mullen
- Invited Talk: Developing Control Plane to Stabilize Hybrid Classical-Quantum Algorithms
- Rohan Basu Roy (Univ. of Utah)
- Invited Talk: Quantum-HPC Supercomputing: From Fault-Tolerant Hardware to a Managed Scientific Ecosystem
- Laura Schulz (Argonne National Lab)
- Invited Talk: Management & Orchestration for Hybrid Quantum-Classical HPC Systems
- Romulo Pinho (Dell Technologies)
Session Chair/Session Host: K. Cain
- STSS: Skill Trust and Signing Service for Secure AI Agent Skill Ecosystems
- Akram Sheriff (Cisco Systems)
- The Integrity of Agentic Intelligence: Tamper-Evident Provenance Receipts for Verifiable Multi-Step Reasoning
- Dalal Alharthi (Univ. of Arizona)
- AI for Scalable Defensive Cyber Log Analysis
- Catherine Schofield (US Air Force), Hayden Jananthan, Jeremy Kepner (MIT Lincoln Laboratory)
- FFUSE: I/O Fault Injection and Resilience Scoring for HPC and AI Workloads
- Anila Ghazanfar (Univ. of Göttingen), Julian Kunkel (Univ. of Göttingen / GDWG)
- High Performance High-Assurance Interval Arithmetic Based Safety Monitoring for Robot Arms on Industrial Multi-Core PLCs
- José Cestero, John Rogan, Franz Franchetti (SPIRAL), Tao Cui (Siemens U.S.)
- Embedded-First Design Principles for Power-Aware High-Performance Computing: A Quantitative Framework
- Amit Jain (Meta)
- Performance Analysis of Open-Source RISC-V Processors for Space Applications
- Richard F. Gibbons III, Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
- Ultra Ethernet Embedded Profile for High Performance Shared Memory Computing
- Kent Dahlgren (Praesum Communications)
Session Chair/Session Host: K. Gettings & I. DeTore
- slurm-qiskit-cluster: A Reproducible Local Testbed for QRMI-Mediated Quantum–HPC Workflows
- Dikran S Meliksetian (Univ. of New Haven)
- Toward Quantum Data Generation for QML Utility
- Kacem Ettahali, Tirthak Patel (Rice Univ.)
- StratiQ: Accelerating Noisy Quantum Simulation with Stratification, Convergence and Checkpointing
- Ashwath Vinodkumar (Vellore Inst. of Tech.), Anamika, Neeraj Goel (IIT Ropar)
- QampLite: Efficient and Scalable Design for Input Embedding in Quantum Machine Learning
- Saber Dinpazhouh, Jason Han, Nick DiBrita, Illya V. Hicks, Tirthak Patel (Rice Univ.)
Session Chair/Session Host: K. Cain
- Hardware-Accelerated Real-Time Data Processing on Heterogeneous FPGA Architectures via AXI4-Lite Interface
- Sudarshan Kale, Vedant Bhandari, Sharda Desai (PuneTech)
- Autonomous Acceleration of Scientific Workloads with Agentic AI
- Thiago Monteiro (UC San Diego), Andrea Guerrieri (HES-SO)
- Reducing Design Time while Developing Energy Efficient FPGA-Based SoCs for Edge AI Inference
- Rishi Agrawal (BITS Pilani), Andrea Guerrieri (HES-SO)
- MANGO: An MTLA-Transformer Accelerator with Reconstructed Normalization and Granularity-Aware Quantization Optimization
- Junhao Zhou, Han Jiao, Yihua Huang (Sun Yat-Sen Univ.)
- 2D Regex Image Processing using FPGA
- Prayag Sridhar (Northeastern Univ.), Ganesh Chennimalai Sankaran, Cong Wang (RENCI), Michael Zink (UMass Amherst), Miriam Leeser (Northeastern Univ.)
- VNS: An Open-Source FPGA SmartNIC for SR-IOV Inter-VM Switching and Function Offload
- Jeffery Lim (MIT Lincoln Laboratory), Martin Herbordt (Boston Univ.)
- SGC-MoE: SmartNIC-Assisted Scheduling for Sparse MoE Training with ZeRO-Offload
- Shining Yang, Anqi Guo, Martin Herbordt (Boston Univ.)
- Autograd-Free Analytic Forces for Machine-Learning Interatomic Potentials on a Forward-Only Tile-Dataflow Accelerator: Method and Measured Limits
- Christian Hanshans and Dominik Kimmerle (Munich Univ. of Applied Sciences), C Herdes (Univ. of Bath)
Session Chair/Session Host: T. Hardjono & S. Pisharody
- Invited Talk: When High Performance Is Not Enough: Rethinking HPC for Enterprise AI
- Pere Monclus (Cisco)
- Invited Talk: TBD
- Ivan Mortimer (GLEIF)
- Invited Talk: Cybersecurity: What Security?
- Yaneer Bar-Yam (NECSI)
Session Chair/Session Host: M. Pelletier
- Inferring Hidden Qubit Geometry in Shared Neutral-Atom Quantum Systems
- Shunyao Mao, Tirthak Patel (Rice Univ.)
- An Entropy-Bounded Thermodynamic Approach for Real-Time Maternal Hemorrhage Monitoring
- Steven D. Harris, Francesca Bonetta-Misteli, Lleyton Martin, Christopher D. Gill, Roger D. Chamberlain, Christine M. O’Brien (Washington University in St. Louis)
- DFD-CR: Decentralized Fluid Dynamics for Congestion Resolution at Multi-Agent Bottlenecks
- Zhenqing Hu and Bin Ren (College of William and Mary)
Session Chair/Session Host: K. Gettings & R. Thoelen III
- DaisyLeak: Timing Side-channel Attacks on Daisy-Chained Thunderbolt
- Claudia Pacori Palomino (Dartmouth Coll.), Junpeng Wan (Purdue Univ.), Jongouk Choi (Univ. of Central Florida), Peter Chin, Kyungtae Kim (Dartmouth Coll.)
- Collaborative Generative AI for Cyber Mission Planning, Analysis, and Reporting
- Aaron Quiroga (MIT AI Accelerator), Jeremy Kepner (MIT LLSC)
- Fingerprinting LLM Inference from eBPF Kernel Telemetry
- Yuanhao Chen, Kevin Limanta, Wuhao Zhang, Kyungtae Kim, Peter Chin (Dartmouth Coll.)
- FLAVIUM: A Multi-Agent System for Reactive Defensive Cyber Operations with Distilled Neural Policies
- Luigi Mastromauro, Yaphet Elias Weldegebriel, Mishel Jyothis Paul, Edwin Kayang, Muslum Ozgur Ozmen, Michel Kinsy (Arizona State Univ.)
- Towards OpenCL in Safety Critical Systems: Lessons Learnt from a RISC-V Space GPU Platform
- Marc Solé i Bonet, Jannis Wolf, Leonidas Kosmidis (Barcelona Supercomputing Ctr.)
- Event-Based Object Detection on Satellite Imagery
- Linus Silbernagel and Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
Session Chair/Session Host: H. Nguyen & M. Pelletier
- CRP-DAR: Condition-Ready Progressive Acceleration for Autoregressive Diffusion Models
- Rui Xu, Han Jiao, Wenjin Huang, Junfeng Li, Chuanle Song, Yihua Huang (Sun Yat-Sen Univ.)
- FPGA-accelerated Semantic 2.5D Mapping for Improved Robot Autonomy on the Edge
- Maria Victoria Gianello, Gaurav Kothamachu Harish, Alireza Ramezani, Miriam Leeser (Northeastern Univ.)
- Design and Analysis of Soft NoC Interconnect Topologies for High-Throughput FPGA Transfers
- James Bickerstaff, Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
- Deterministic Dual-Path Inference on Resource-Constrained FPGAs
- Syed Daniyal Naqvi (University of Liverpool)
- Low Latency FPGA Implementation of a Neural Network-based Particle Physics Trigger
- Charles Lange (Boston Univ.), Dana Diaconu, Miriam Leeser (Northeastern Univ.)
- SpecReuse: Spectral Graph Reuse for Efficient Vision GNN Inference on FPGAs
- Isabella Bernhardt Eiliya, Anvitha Ramachandran, Dhruv Parikh, Viktor Prasanna (USC)
- Numerical Kernels on a Spatial Accelerator: A Study of Tenstorrent Wormhole
- Maya Taylor (Univ. of Illinois Urbana-Champaign), Carl Pearson, Luc Berger-Vergiat (SNL), Giovanni Long (UC Santa Barbara), Jan Ciesko (SNL)
Session Chair/Session Host: J. Kepner & A. Reuther
- Keynote Talk: 30 Years of HPC==The Blueprint of AI
- Gregory Kurtzer (CIQ)
Session Chair/Session Host: B. Raut & L. Zaidenberg
- Decoupled Azimuth–Elevation AoA Estimation Exploiting Kronecker-Separable Steering Matrices
- Faizan A. Khattak (University of Leeds), Ian K. Proudler, Stephan Weiss (Univ. of Strathclyde), Fazal-E Asim (Fed. Univ. Ceara Fortaleza)
- Full-Stack Benchmarks of a Blackwell-Generation AI Cluster: Compute, Interconnect, and End-to-End Training
- Shaohao Chen, Lauren Milechin, Christopher N. Hill (MIT)
- Application Energy vs System Energy: A Multi-Workload Evaluation in FPGA-Accelerated HPC Nodes
- Paolo Palazzari, Marco Faltelli, Francesco Iannone (ENEA), Pasquale Tommasino (Sapienza Univ.)
- Walls Of Computing
- Sharwari Bhosale, Zhenqing Hu, Christopher Scrosati, Aniruddha Bhattacharjee (SLAC National Accelerator Laboratory), Sadasivan Shankar (SLAC and Stanford Univ.)
- OpenDwarfs 2.0: Evolving the Berkeley Dwarfs in OpenCL and OpenMP [Best Student Paper Award]
- Nabayan Chaudhury and Wu-chun Feng (Virginia Tech)
Session Chair/Session Host: K. Keville
- A Scalable, Communication-Efficient Distributed Algorithm for Exact Square Counting
- Shubhashish Kar, Shaikh Arifuzzaman (UNLV)
- Parallel Canonical Labelling for Graph Isomorphism Testing
- Jim Haslett, Daniel Grosu (Wayne State Univ.)
- NI-ORCA: Parallelising the Counting of Orbits of Non-Induced Graphlets
- Syed Ibtisam Tauhidi (Queen’s Univ. Belfast), Arindam Karmakar (Tezpur Univ.), Thai Son Mai, Hans Vandierendonck (Queen’s Univ. Belfast)
- Fast Analytic Eigenvector Extraction and Partial Analytic Eigenvalue Decomposition
- Faizan A. Khattak (University of Leeds), Mohammed Bakhit, Ian K. Proudler, Stephan Weiss (Univ. of Strathclyde)
- Benchmarking GNN Inference on the Intel Core Ultra NPU: A Latency, Quantization, and Energy Analysis
- Yusuf Talha ARABACI, Emrullah DEMİRAL, Ömer Faruk ACAR (Karabuk Univ.)
- CompJouleS: A Cross-Platform Energy Estimation Tool for Machine Learning Algorithms on Heterogeneous Computing Architectures
- Aniruddha Bhattacharjee (SLAC National Accelerator Laboratory), Murat Isik (Purdue Univ.), Vedant Karia (Univ. of Texas at San Antonio), Jens Pedersen (Tech. Univ. Denmark), Sadasivan Shankar (SLAC and Stanford Univ.)
- A Quantitative Evaluation of Neuromorphic and GPU Architectures for High-Performance Edge Computing
- Mark Barnell, Courtney Raymond, Lisa Loomis, David Wise (AFRL), Darrek Isereau, Daniel Brown, Francesca Vidal (SRC)
- Simplify to Amplify: Achieving Information-Theoretic Bounds with Fewer Steps in Spectral Community Detection
- Sie Hendrata Dharmawan, Peter Chin (Dartmouth Coll.)
- Scalable Smart Grid Attack Dataset Generation for IDS Using Containerized NATIG Co-Simulation
- Kenneth Watts (UMass Lowell)
Organizers: J. Mullen, H. Jananthan, R. Thoelen III
- Invited Talk: Teaching HPC Through Local LLM-RAG Pipelines: Lessons from Hands-On Workshops
- Sam Corey (MIT ORCD)
- Invited Talk: Equipping HPC Practitioners for Container Platform Operations
- Robert Thoelen III (Pratt & Whitney)
- Invited Talk: AI Can Write the Code. Now What? Rethinking RSE Education
- Sandra Gesing (US-RSE)
Session Chair/Session Host: N. Pitsianis
- Scalable Cross-Aircraft Flight Phase Classification Using Transfer Learning and ADS-B Data [Outstanding Paper Award]
- Jacob Kiefer (US Air Force) and Sheila Alemany (MIT Lincoln Laboratory)
- Are Multidimensional Models Worth it in Demand Forecasting?
- Cheng-Jui Fan (Shanghai High School), Nikolay Aristov, Elenna R. Dugundji (MIT)
- Learning to Select Sparse Linear Solvers with Convolutional Neural Networks
- Artemis Pados, Alan Edelman, Emmanuel Lujan, Daniel Pickard, Felipe Tome, Christopher Rackauckas (MIT)
- An LLM-Based Triage System for GPU Numerical Failures in High-Performance ML Software
- Aryan Shah (Univ. of North Texas)
- Architecture and Performance Tradeoffs with Small Vision Transformers for Image Processing on the Edge
- Ian Peitzsch, Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
- HEALTH-RAG: A Deployment-Oriented Retrieval-Augmented Framework for Grounded Healthcare Question Answering
- Jamal Haider Rizvi, Wali Mohammad Abdullah (Concordia Univ. of Edmonton)
Session Chair/Session Host: D. Dixit & C. O’Joy
- Easy Multi-Precision Acceleration via Vectorization
- Chansup Byun, Piotr Luszczek, Jeremy Kepner, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Hayden Jananthan, Michael Jones, Lauren Milechin, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Peter Michaleas (MIT)
- Performance Evaluation of GPU-Based Random Number Generators in MARLEY [Outstanding Student Paper Award]
- Yohannes Abateneh (Florida A&M Univ.), Steven Gardiner (Fermi National Accelerator Laboratory), Kimieka Dunkley, Hongmei Chi (Florida A&M Univ.)
- Accelerating Spiking Neural Network Inference with RISC-V Vector Extension [Outstanding Student Paper Award]
- Myles Fernau, Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
- LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs [Outstanding Paper Award]
- P. Kuppili (Univ. of Utah), Y. Qian, S. Handagala (Northeastern Univ.), M. Zink (UMass Amherst), M. Leeser (Northeastern Univ.), R. Ricci (Univ. of Utah)
- Performance and Portability of a Production Fluid–Structure Interaction Solver Across NVIDIA GPU Generations [Outstanding Paper Award]
- Ella Fortenbery, Jorik Stoop, Ayman Yousef, Amanda Randles (Duke Univ.)
Session Chair/Session Host: J. Mullen
- MLOps Framework for Deploying LLM-Based Climate-Adjusted Portfolio Optimization in Banking
- Milan Parikh (Cytel), Rohit Nimmala, Viswanathan Ranganathan, Jagrut Nimmala (Independent Researcher)
- Oscillatory Graph Preprocessing: A Reusable, Amortizable Structural Prior for Deep GCNs
- Fernando Vera Buschmann, Isidro Gauto, Dahlia Musa, Horacio G. Rotstein, Vincent Oria (New Jersey Inst. of Tech.)
Session Chair/Session Host: N. Pitsianis & R. Thoelen III
- Optimizing Graph I/O for In-memory Computations
- Ali Brooks, George M. Slota (RPI)
- LOOM: A Team-of-Agents Architecture for High-Performance Graph Analytics Workflows with Arkouda and Arachne
- Asha Saxena, David A. Bader (New Jersey Inst. of Tech.)
- Scaling Exact Substructure Extraction Beyond the Five-Vertex Graphlet Wall
- Mohammad Dindoost, Bartosz Bryg, David A. Bader (New Jersey Inst. of Tech.)
- Topology Survives Shuffling: Robust Long-Context Memory Access Prediction for Graph Analytics [Best Student Paper Award]
- Dongyan Sun, Neelesh Gupta (USC), Rajgopal Kannan (DEVCOM Army Research Lab), Viktor Prasanna (USC)
- ML-Guided Parameter Configuration Selection for Top-Down Stochastic Block Partitioning [Outstanding Student Paper Award]
- Saikat Dey, Wu-chun Feng (Virginia Tech)
Organizers: T. Mattson, B. Brock & S. McMillan
Session Chair/Session Host: J. Kepner & A. Reuther
- Keynote Talk: GPU-Initiated Data Access: A Disruption to the Storage Industry
- CJ Newburn (Nvidia)
Session Chair/Session Host: A. Wright & J. Mullen
- Multi-Scale AI Training for Satellite Imagery
- Victor M. Vergara (AeroVironment/AFRL), Amanda Fetzer, Evan T. Kain (AFRL)
- RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
- Eugene Hauptmann, Nataliya Kosmyna (MIT)
- PipeDiff: A ReRAM-Based Pipelined Compute-in-Memory Accelerator for Diffusion Model Inference
- Arjuman Ara Mimi, Su-in Yi (Texas A&M Univ.)
- Transition-Aware Backend Dispatch for Edge LLM Inference
- Alaaddin Goktug Ayar, Martin Margala (Univ. of Louisiana at Lafayette)
- PoolStateLens: Measurement-Grounded Evaluation of an On-Switch Memory Pool for LLM Training
- Xiteng Yao, Martin Herbordt (Boston Univ.)
Organizers: F. Franchetti & N. Zhang
- Spiral Tutorial – https://www.spiral.net/tutorial-spiral.html
Session Chair/Session Host: N. Pitsianis & S. Mehta
- How Far Can GPGPU Push the OFDM Lower-PHY? A MIMO, Bandwidth, and Numerology Study on GB10
- Jordan Vrtanoski, Alexandre Loureiro (Connect 5G)
- CREDIT: Cost-guided Reduction-reuse with Efficient DSMEM Inter-CTA Tiling [Outstanding Paper Award]
- Zhengxiong Li, Tsung-Wei Huang, Umit Ogras (Univ. of Wisconsin)
- GPU-Accelerated Branch-and-Reduce for Feedback Vertex Set
- Ashwina Kumar, Rupesh Nasre (Indian Inst. of Tech. Madras)
- DPCP: A Python Runtime Layer for Portable GPU Kernel Execution
- Niteya Shah, Wu-chun Feng (Virginia Tech)
Session Chair/Session Host: K. Keville
- Mixed-Precision Restarted Anderson Acceleration: Attainable Accuracy at the Working-Precision Limit
- Stephen J. Thomas (Lehigh Univ.)
- Local Forces, Global Structure: A Unified View of t-SNE, UMAP, and Their Successors
- Constantine Roshi and Dimitris Manolakis (MIT Lincoln Laboratory)
- Less is More: Targeted Mamba Deep Learning Integration for Low-Dose CT Denoising
- Changze Li, Junkai Zhang, Wu-chun Feng (Virginia Tech)
- Robust IoT Security: A Holistic Evaluation of Machine Learning-Based Intrusion Detection
- Shaibal Das, Fairuz Nawar, Hao Yin, and Abu Asaduzzaman (Wichita State Univ.)
- Fast Recovery Is Not Enough: Correct Goodput and Trajectory Fidelity for Foundation-Model Training on Preemptible Clusters
- Naeem Khoshnevis, Yasin Mazloumi, Camilo Brown-Pinilla, Timothy Ngotiaoco, D. Balamurugan, Sarah Leinicke, and Max Shad (Harvard Univ.)
- Subject-Independent Muscle Fatigue Detection from Surface EMG with a Spiking Neural Network and Verified FPGA RTL
- Yusuf Kerim Kaymakçı (Baykar Science HS), Zeynep Bahat (Ordu Bahçeşhir Science and Technology High School), Zeynep Elveren (Deutsche Schule Istanbul), İsmail Can Dikmen (İstinye Univ.)
Session Chair/Session Host: P. Luszczek & S. Mehta
- Optimizing Very Large Integer Multiplication Using GPUs
- Aarushi Aggrwal, Nikita Borisov (Univ. of Illinois Urbana-Champaign)
- Autotuning GPU Thread Block Sizes for Large-Scale Stencil Applications Using the OPS DSL
- Kevin Antonio Nava Garcia, Istvan Zoltan Reguly (Pázmány Péter Catholic University Faculty of Information Technology and Bionics)
- GCStack+ and GCScaler+: Closing the Calibration Gap in GPU Performance Modelling for Modern Microarchitectures
- Tim Lühnen, Ulf Kulau (Hamburg Univ. of Tech.), Devashree Tripathy (IIT Bhubaneswar), Sohan Lal (Tech. Univ. Berlin)
- HARP-ME: Closure-Driven Exact Induced Motif Enumeration on GPUs
- Ashwina Kumar, Rupesh Nasre (Indian Inst. of Tech. Madras)
- The Two Faces of Abstraction Regret: Control-Flow and Memory-Layout Limits of GPU DSLs on Irregular Automata
- Alessandro Potenza (Politecnico di Milano)
Session Chair/Session Host: S. Gomez
- Mission Computing: Bounded Runtime Orchestration for Heterogeneous FPGA-Based Spacecraft Computing
- Nitesh Bakhati, Paolo D’Alessandro (University of Florida), Dave Ojika (Flapmax), Aidan Osuch (UCLA), Ben Cragin (USC)
- Adaptive Trust-Aware Byzantine-Resilient Federated Learning for Zero-Trust Tactical Edge AI Systems Under Adversarial Conditions
- Philip W. Chan (UMGC), Daniel Ku (US Army)
- On-board ML for Trace Gas detection in Imaging Spectroscopy data
- Vít Růžička, Adam Chlus, Andrew Thorpe and David Thompson (NASA-JPL)
- Depthwise vs. Channel-Mixing Residual Blocks: Hardware and Accuracy Evaluation
- Ryan Gaffere (Johns Hopkins Univ.)
- Projection-Aware Approximation for Efficient AI Accelerators
- Leonard MacEachern, Jonathan Levine (Carleton Univ.)
- Which Agentic AI Architecture for Space Security?
- Youssef Ejjiyar, Marc Lacoste (Orange)
- When FLOPs Mislead: Benchmarking Neural Network Inference Energy on Apple Silicon
- Aryan Shah, Romir Kadiam (Univ. of North Texas)
- Smart Drones: SWAP-C Analysis for Agentic Drone Operations
- Holt Russell, Brian Wheelhouse, Thomas Russell, Timothy Kokotajlo, Dylan Carpenter, Dylan Che, Joanna Russell, Benjamin Spellman (US Air Force), and Lei Hamilton (MIT Lincoln Laboratory)
- Performance-Energy Characterization of KV-Cache Quantization for Edge LLM Inference on Apple Silicon
- Khush Patel, Tanmay Sharma, Manuel Mazzara (Innopolis Univ.)
- Proactive Security for Large-Scale AI Inference: Adaptive Container Orchestration in Kubernetes
- Akram Sheriff (Cisco Systems), Zsolt Németh (R6 Security), Ken Huang (Distributedapps.ai)
Session Chair/Session Host: M. Pelletier & N. Zhang
- Cross-Model Cross-Language AI Coding Agent Performance: Accuracy and Speed of Parallel CLRS Algorithms
- Shiqi Cheng, Evelyne Ringoot, Rabab Alomairy, Alan Edelman (MIT)
- ACBench: A Benchmark Suite for Auditing Retry and Profiling Feedback in LLM-Driven Code Optimization
- Xiteng Yao, Ziyuan Chu, Feng Tai, Bangjie Xue, Yigong Hu, Martin Herbordt (Boston Univ.)
- Performance Implications of CPU-Assisted Embedding in GPU-Based RAG Systems
- Mona Minakshi, Shamima Najnin (Intel)
- SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling [Best Paper Award]
- Alaina Kolli (MIT), Theodoros Xenakis (MIT & NTNU), Utkarsh Utkarsh, Pengfei Cai, Rafael G´omez-Bombarelli, Alan Edelman, Christopher V. Rackauckas (MIT)
- Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers
- Menachem Finkelstein, Diana Levy, Zohar Yakhini, Sarel Cohen (Reichman Univ.)
- CLAS: A Unified Static-and-Runtime Loop Analysis Framework for Compiler Optimizations
- Mark Samuel, Samruddhi Dhakulkar, Prachi Pandey, Haribabu P, S D Sudarsan (C-DAC)
Organizers: J. Kepner & A. Reuther
- Invited Talk: Temporal Block Partition Graph Challenge
- Shahar Somin (MIT)
- MPI-Leiden: Fully Distributed Community Detection for Very Large Graphs
- Pavel Serhiayenka, Alan D. George (Univ. of Pittsburgh / NSF SHREC Center)
- An Empirical Study of Kernel-Aware Graph Sparsification for Parallel BFS and SSSP
- Rakibul Hassan, Shaikh Arifuzzaman (UNLV)
- An Equal-Footing Comparison of Distributed Strongly Connected Component Algorithms
- Preston Piercey, Trevor Steil, Roger Pearce (LLNL)
- DUET: Efficient Subgraph Matching on a Practical Heterogeneous GPU-PIM Platform
- Yiheng Yang, Yi Zhang, Yu Huang, Deting Chen, Qihang Qiu, Long Zheng, Xiaofei Liao, Hai Jin (Huazhong Univ. of Science and Tech.)
- Triangle-Sparse and Small-Diameter Networks: A Standing Challenge for Graph Laplacian Solvers
- Dimitris Floros (Duke Univ.), Nikos Pitsianis (Aristotle Univ. of Thessaloniki), Xiaobai Sun (Duke Univ.)
- GetterTri: A GPU Approach to Accelerate the Triangle Counting Computation
- Nicola Basciu, Lorenzo Cardone, Stefano Quer (Politecnico di Torino)
- Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) Arrays
- Jeremy Kepner, Hayden Jananthan, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Chansup Byun, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Piotr Luszczek, Peter Michaleas, Lauren Milechin, Chasen Milner, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Alex Pentland (MIT Lincoln Laboratory)
- The Anonymized Network Sensing Graph Challenge in PostgreSQL via OneSparse and GraphBLAS
- Michel Pelletier (MIT), Timothy A. Davis (Texas A&M Univ.), Timothy G. Mattson (Univ. of Bristol)
- FOAMS: Efficient Anonymized Network Sensing using Compact Matrix Format and Multi-Stage Pipeline Parallelism
- Jun Mai, Qinggang Wang (Huazhong Univ. of Science and Tech.), Haonan Wu (Boston Univ.), Pengcheng Yao, Yu Huang, Long Zheng, Xiaofei Liao, Hai Jin (Huazhong Univ. of Science and Tech.)
- Processing Network Sensing Data with the Gmap Graph System
- Daniel Osei, Chanaka Hettige, Martin Swany (Indiana Univ.)
Session Chair/Session Host: J.Ghanem & A. Reuther
- Adaptive Row Selection Meets Asynchrony in Randomized Kaczmarz
- Evan Coleman (Univ. of Mary Washington)
- Increasing Metadata Parallelism through Hybrid Inode Ownership
- Matthew Curtis-Maury, Jian Hu, Bulli Venkata Rajesh Vipperla, Sushrut Bhowmik, Grey Files, Bryan Prosser (NetApp)
- Single-Source Scalar and SIMD Particle Kernels with Compile-Time Memory Abstractions
- Julian Deller-Yee, Manish Kumar Mishra, and Hans Joachim Bungartz (Tech. Univ. Munich)
- Diff-GNN: Differentiable Graph Neural Networks for Joint Hardware-Software Partitioning and Scheduling
- Siddhartha Shankar Das (PNNL), Zakaria Mehrab (Univ. of Virginia), James Kotary, Rounak Meyur, S. M. Ferdous, Erdal Mutlu, and Mahantesh Halappanavar (PNNL)
- Keynote Talk: Intelligence Where the Task Is
- Daniela Rus (MIT)
- Parametrized Exact Hardware-Software Partitioning
- Cameron Ibrahim, S.M. Ferdous, Erdal Mutlu (PNNL), Ilya Safro (Univ. of Delaware) and Mahantesh Halappanavar (PNNL)
- Joint Optimization of HW/SW Partitioning and Scheduling with Policy Gradient Methods
- James Kotary (PNNL), Zakaria Mehrab (Univ. of Virginia), Siddhartha Shankar Das, Rounak Meyur, S. M. Ferdous, Erdal Mutlu, and Mahantesh Halappanavar (PNNL)
Session Chair/Session Host: K. Keville
- CARP-RAG: Compute-Aware Adaptive Retrieval for Efficient Retrieval-Augmented Generation
- Abhishek Kumar (Univ. of Cumberlands)
- Why Runtime-Only Reinforcement Learning Struggles with Compiler Optimization on Blackwell GPUs
- Xiteng Yao, Martin Herbordt (Boston Univ.)
- From 8 Seconds to 370 ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
- Mohamed Amine Bergach (Illumina)
- GPU-Based Parallelization of Differential Evolution Variants: Application to Solving Large Systems of Nonlinear Equations
- Diogo Martins, Bruno Silva, Luiz Guerreiro Lopes (Univ. of Madeira)
- Performance Engineering of Retrieval Architectures for Enterprise Supply-Chain Question Answering
- Alex Carroll, Elenna Dugundji (MIT)
Session Chair/Session Host: X. Sun & J.Ghanem
- Invited Talk: Compiler 2.0: Compilers in the Era of Machine Learning
- Saman Amarasinghe (MIT)
- Bounding Both Ends of an Anticipatory FPGA Scheduler
- Syed Daniyal Naqvi (University of Liverpool)
- SSTA-Trace: LLM-Learned Scheduling of Statistical Static Timing Graphs on GPU
- Chedi Morchdi (Texas A&M Univ.), Cheng-Hsiang Chiu (Univ. of Wisconsin), Yi Zhou (Texas A&M Univ.), Tsung-Wei Huang (Univ. of Wisconsin)
- HLS4LFADS: A Framework for Real-Time and Parallel Neural Decoding on FPGAs
- Chi-Jui Chen (National Yang-Ming Chiao Tung Univ.), Yan-Lun Huang (Univ. of Texas Austin), Xiao-Han Liu, Atharva Mattam, Hao Fang (Univ. of Washington), Ling-Chi Yang, Chia-En Chang (National Yang-Ming Chiao Tung Univ.), Elham E. Khoda, Eli Shlizerman, Amy L. Orsborn, Scott Hauck, Shih-Chieh Hsu (Univ. of Washington), and Bo-Cheng Lai (National Yang-Ming Chiao Tung Univ.)
- Agentic CRAMP: Runtime-Adaptive Scheduling for Scalable Parallelism of Classifiers and Regressors on Distributed and Multicore Systems
- Baidya Nath Saha, Wali Mohammad Abdullah, Md. Morshedul Islam (Concordia Univ. of Edmonton)
Session Chair/Session Host: K. Keville
- Linux Support for CPU Stress Test on Large-Scale AMD EPYC Systems
- S. Biplab Raut (AMD)
- Intent-Preserving FPGA-SoC Integration Using a Design-Time Metadata Intermediate Representation
- Lewis McLaughlin, Louise H. Crockett, Robert W. Stewart (Univ. of Strathclyde)
- Satoorn: A Source-to-Source Conversion Tool for Safety Critical Model Based Generated GPU Code
- Marcos Rodriguez, Alejandro J. Calderón (Ikerlan Research Center), Leonidas Kosmidis (Barcelona Supercomputing Ctr.), Irune Yarza (Ikerlan Research Center)
- GRAPHSCHED: Heterogeneous Graph-Augmented Multi-Objective Reinforcement Learning for Energy-Aware HPC Workload Scheduling
- Kyrian Adimora, Hongyang Sun (Univ. of Kansas)
- LockRoute: A Spatial Locking Framework for Parallel Global Routing
- Amber Thrall (Washington State University), Vidya Chhabria (Arizona State Univ.), S M Ferdous, Mahantesh Halappanavar (PNNL), Bala Krishnamoorthy (Washington State University)
- ML Based Modeling of XSEDE/ACCESS Usage and Performance data for HPC Job Wait Time Estimates and Resource Allocation
- Abani Patra, Bipin Gaikwad (Tufts Univ.), Nikolay Simakov, Joseph White, Thomas Furlani (Univ. of Buffalo)
- Breaking the Interconnect Bottleneck in LLM Serving via Congestion-Aware Scheduling over Heterogeneous Interconnects
- Dong Liu (UCLA), Yanxuan Yu (Columbia Univ.), Chang Liu (Univ. of Illinois Urbana-Champaign), Eric Jiang (UCLA), Tony Geng (Rice Univ.), Ying Nian Wu (UCLA)
- TriShield: A Unified ASIC Architecture for Composable TEE, FHE, and Differential Privacy with Leakage-Aware Formal Guarantees
- Akram Sheriff (Carnegie Mellon Univ.)
Session Chair/Session Host: N. Zhang & P. Luszczek
- BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training
- Bigyan Ghimire, Jon Calhoun (Clemson Univ.)
- To CHAI or Not To CHAI: Chapel+AI for Training and Inference
- Mohamed Abdelnaby, Paul Sathre, Wu-chun Feng (Virginia Tech)
- Selector–Solver Coordination in Vision Language Model-Guided Combinatorial Optimization
- Yadu Shaji Nair, Elenna Dugundji (MIT)
- Software-Customized LLM Serving on Large-HBM GPUs
- Zhihui Du, Diptorup Deb, Clint Greene, Yao Liu, Rishi Madduri, Mukhil Azhagan Mallaiyan Sathiaseelan, Debasis Mandal, Phani Vaddadi, Viswanath Vadlamani, Anuya Welling (AMD)
Session Chair/Session Host: M. Pelletier
- Traffic Matrices for Anonymized High Performance Analysis of Physical/Data/Network/Transport/ Session/Presentation/Application Layers
- Martin Satterfield, Justin Wynn, Grant Blakely, Tyler House, DJ Hovermale (UAH), Hayden Jananthan, Jeremy Kepner (MIT)
- Clique search in simultaneously partitioned graphs
- Peter Schmidt, Sandor Szabo, Bogdan Zavalnij (Renyi Institute)
Organizers: B. Hoppe & T. Kaler
- Invited Talk: Beyond Correct — Teaching AI to Actually Solve Hard Problems
- Alvin Cheung (Berkeley)
- Invited Talk: The Post-Developer Era
- Tim Kraska (MIT)
- Invited Talk: LLMs are Great at Code Generation, but What About Code Adaptation?
- Justin “Goju” Gottschlich (Merly)
- Invited Talk: Green Means Stop – Building the Reward Signal Code Optimization We Never Had
- Jatin Ganhotra (IBM)
Organizers: D. Burrill, V. Gadepally & C. Prothman
- Invited Talk: Intel Crescent Island
- Hong Jiang (Intel)
- Invited Talk: Rearchitecting the Datacenter Lifecycle for AI
- Chaojie Zhang (Microsoft)
- Invited Talk: Language Model Architecture
- Roger Waleffe (Nvidia)
- Invited Talk: Powering AI: Energy, Control and Consequences
- Shashank Yellapantula (NLR)
Session Chair/Session Host: J. Kepner & A. Reuther
- Keynote Talk: HPEC 30th Year Anniversary: Reflections and a Forward-Looking Perspective
- Dave Martinez (MIT LL)
Organizers: P. Luszczek & K. Keville
- Invited Talk: Recent Progress of Ozaki Schemes
- Katsuhisa Ozaki (Shibaura Inst. of Tech.)
- Invited Talk: Variable Precision Formats that Preserve Dynamic Range
- John Gustafson (Arizona State Univ.)
- Invited Talk: Mixed model, mixed method, and mixed precision Runge–Kutta method
- Sigal Gottlieb (UMass Dartmouth)
Session Chair/Session Host: L. Anderson
- A Hybrid MPI+Multi-GPU Solver for Coupled Voltage-Calcium Dynamics in Neuronal Networks
- Zachary M. Miksis, Gillian Queisser (Temple Univ.)
- ParAGraph: Parallel Agent-Based Graph Modeling
- Jayanta Mukherjee (Purdue Univ.), Minhyuk Park, George Chacko, Tandy Warnow (Univ. of Illinois Urbana-Champaign), Ananth Grama (Purdue Univ.)
- On Mixed-Precision Iterative Methods for Nearly Decomposable Principal Eigenvector Problems
- Vasileios Kalantzis, Mark Squillante, Chai Wah Wu (IBM Research)
- An Analysis of Statistical Techniques to Evaluate MPI Tuning Parameters
- Shannon Kinkead (SNL), Amanda Bienz (University of New Mexico), Whit Schonbein, Matthew G. F. Dosanjh (SNL)
- Efficient High-Precision Floating-Point Arithmetic Using Integer Units: A First Look
- Yunhao Lan, Larry Tang, Naifeng Zhang, Franz Franchetti (Carnegie Mellon Univ.)
- Optimizing Quantized GEMV for LLM Inference
- Vishva Ranjan Singh (IIT Bhubaneswar), Shivam Godayal, Sanyam Garg (IIT Guwahati), Devashree Tripathy (IIT Bhubaneswar), Aryabartta Sahu (IIT Guwahati)
- Simulation and Performance Study of Mixed HPC Application Workloads Using NVIDIA Multi-Instance GPU Technology: Cost, Performance, and Energy Efficiency Analysis Across A100, H100, H200, and B200 GPUs
- Hsingbung Chen (Frisco HPC Performance engineering group)
- Sema: Embedding Go in Python for High-Performance Computing
- Mahnoor Syeda (Punjab Colleges), Khalil E. A. Abdulgawad (Istanbul Aydin Univ.)
- Build-Time-Specialized SAR Image Formation from x86 Ground Segments to Mobile ARM Platforms
- Maron Schlemon (German Aerospace Center), Martin Schulz (Tech. Univ. Munich), Rolf Scheiber (German Aerospace Center)
- GPU-Accelerated Processing of Weighted Top-k Dominating Queries over Incomplete Dataset
- Nabil Faiyaz Sadi, K. M. Azharul Hasan (Khulna Univ. of Engr. and Tech.)
- High-Performance Extended Precision Integer Arithmetic with Intel(R) Advanced Matrix Extensions (Intel(R) AMX)
- Ted Painter (Intel)
- Offline vs Real-Time Validation and Performance Assessment of Heimdall Detection Pipeline for Fast Radio Bursts
- Valentina Cesare, Giovanni Naldi, Francesco Fiori (INAF-IRA), Andrea Geminardi (IUSS Pavia, Univ. of Trento, INAF-OAC), Adrian De Barro, Alessio Magro, Hayley Camilleri (ISSA, Univ. of Malta)
Session Chair/Session Host: B. Sroka & H. Nguyen
- Numerically Stable Cholesky-QR on GPU via Mixed-Precision Randomized Preconditioning [Outstanding Student Paper Award]
- James E. Garrison, Chao Chen, Ilse C. F. Ipsen (North Carolina State Univ.)
- Simulation of Custom-Precision OCP MX Block Floating-Point Formats and Arithmetic
- Maliha Islam and Mantas Mikaitis (University of Leeds)
- Low/Mixed-Precision for Level 1 BLAS, a Test Case: AXPY
- Pedro Valero-Lara, Maria Patrou, Piyush Sao, Keita Teranishi, Jeffrey S. Vetter, Narasinga Rao Miniskar, Oscar Hernandez, Sudip Seal (ORNL)
- Closing the Null Space: Guidance-Aware Quantization for Classifier-Free Diffusion
- Abdullah Al Shafi, Sumaiya Rahim Suma (Khulna Univ. of Engr. and Tech.)
- Introducing the MIT Processor Database
- Emanuele Del Sozzo (MIT)
Session Chair/Session Host: TBD
Session Chair/Session Host: B. Sroka & P. Luszczek
- Structural-Similarity-Preserving Lossy Data Compression for CPUs and GPUs
- Alex Fallin, Martin Burtscher (Texas State Univ.)
- The MIT AI Systems Dataset: Environmental, Compute, and Workload Traces
- Piotr Luszczek, Daniel Burrill, William Bergeron, Vijay Gadepally, Matthew Hubbell, LaToya Anderson, William Arcand, David Bestor, Alexander Bonn, Chansup Byun, Michael Houle, Michael Jones, Peter Michaleas, Julia Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, and Jeremy Kepner (MIT Lincoln Laboratory)
- Augmented Geometric Multi-Resolution Analysis: A Fixed-Length, Parallelizable Multi-Scale Feature Extractor
- Minh N. Bui, Kevin Limanta, Peter Chin (Dartmouth Coll.)
- Arnoldi Iteration Based Tensor Rank Estimation
- Matthew D. Merris, Tim Andersen (Boise State Univ.)
- Creating SquashFS Images to Support HPC and AI Datasets
- Julie Mullen, Albert Reuther, LaToya Anderson, Michael Jones, William Arcand, William Bergeron, David Bestor, Alex Bonn, Daniel Burrill, Chansup Byun, Vijay Gadepally, Michael Houle, Matthew Hubbell, Hayden Jananthan, Piotr Luszczek, Peter Michaleas, Guillermo Morales, Andrew Prout, Antonio Rosa, Chuck Yee, and Jeremy Kepner (MIT Lincoln Laboratory Supercomputing Center)
Session Chair/Session Host: M. Pelletier
- A Data-Driven Analysis of Customer Churn Prediction Using Machine Learning and Business Intelligence Techniques
- Nikith Somisetty, Alvin Chin (Univ. of Illinois Chicago)
Session Chair/Session Host: C. Roberts & L. Zaidenberg
- Portable Extended-Precision Floating-Point Arithmetic for Scientific Computing Kernels
- Manas Vishal (UMass Dartmouth)
- Physically Reconfigurable High-Performance Computing with Volume Mount Devices
- Dimitar Dimitrov, Erik Strand, Neil Gershenfeld (MIT Center for Bits and Atoms)
- High performance scalable Ising optimization with GraphBLAS
- Giovanni Gaio, Denis Jelovina, Albert-Jan Yzelman (Huawei Technologies)
- Transparent NumPy Acceleration with SVE on Fugaku
- Kuldeep Pal, Deepika H. V., Hari Babu P., S. A. Kumar and S. D. Sudarsan (C-DAC)
- Sensitivity-Driven Optimization of Sparse Aggregation for GNN Message Passing on GPUs
- Zhihui Du, Yao Liu, Geoffrey C. Martin-Noble, Tres Popp, Mukhil Azhagan Mallaiyan Sathiaseelan, James E. T. Smith, Phani Vaddadi, Viswanath Vadlamani, Anuya Welling (AMD)
- Performance Evaluation of Stabilized Corrections for Mixed Precision Runge–Kutta Methods
- César Herrera (Purdue Univ.), John Driscoll, Sigal Gottlieb, Zachary J. Grant, Tej Sai Kakumanu (UMass Dartmouth), Andrew Christlieb (Michigan State Univ.)
Session Chair/Session Host: C. Roberts & M. Pelletier
- An All-Atom Molecular Dynamics Workflow for the Realistic Integration and Simulation of Bolalipids
- Eirini Evangelinos, Isaac Tucker (Broad Institute of MIT and Harvard), Giovanni Traverso (Broad Institute of MIT and Harvard, Brigham and Women’s Hospital)
- NOMAD: A Hybrid Power-Capping Framework for Heterogeneous Compute Nodes
- Vignesh Adhinarayanan, Wu-chun Feng (Virginia Tech)
- Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming
- Hao Mao (Hong Kong Polytechnic Univ.), Xu T. Liu (Univ. of Washington, AWS), Shuai Lu (Jiangxi Univ. of Finance and Economics), Peng Zhao (EEO Education Technology), Wenzhen Jiang (Chongqing Medical Univ.), Yuntian Chen (Eastern Inst. of Tech., Ningbo)