Days: Wednesday, August 26th Thursday, August 27th Friday, August 28th
View this program: with abstractssession overviewtalk overview
Designing for Trust, Transparency, and Efficiency in Scientific Computing
Scientific computing is entering a regime where nondeterminism, opaque AI decisions, and energy costs are no longer edge cases—they define everyday practice at scale. In this keynote, I will present practical methods that make modern scientific workflows more trustworthy, transparent, and efficient. I will demonstrate how graph-based analysis can pinpoint the sources of nondeterminism in large HPC simulations; how fine-grained provenance can make AI-driven workflows auditable and easier to explain; and how predictive engines can avoid redundant computation in neural architecture search, cutting both runtime and energy. I will conclude with a broader vision: scientific ecosystems that intentionally couple data, experiments, and computation—so results are not only faster, but also reproducible, explainable, and ready for reliable reuse.
| 11:00 | Velvet: Parallel Divide-and-Conquer in Safe Rust (abstract) |
| 11:25 | Communication Offloading on SmartNIC DPUs: A Quantitative Approach (abstract) |
| 11:50 | MemBridge: Bridging the Static-Dynamic Semantic Gap in Memory Profiling via Variable-Centric Instrumentation (abstract) PRESENTER: Feng Wentao |
| 12:15 | TBF: A Tunable Blocking-and-Fusion Algorithm for Efficient GPU Symmetric Rank-2K Updates (abstract) PRESENTER: Lijuan Hu |
| 11:00 | MoProteus: LLM-Driven Multi-Version Operator Generation for Energy-Aware Scheduling in Heterogeneous Cloud-Edge Environments (abstract) PRESENTER: Yiwen Liu |
| 11:25 | Sufficiency in Data Centers: Energy Aware Resource Recommendation System (abstract) |
| 11:50 | Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows (abstract) |
| 12:15 | Machine Learning-Enhanced Resource Optimization for Bioinformatics Workflows in the Cloud (abstract) |
| 11:00 | Characterizing Needleman-Wunsch Sequence Alignment on the Graphcore IPU: Memory Trade-offs and Optimization Strategies (abstract) |
| 11:25 | Backend Code Generation for Graph DSLs Targeting Diverse Accelerator Platforms (abstract) PRESENTER: Ashwina Kumar |
| 11:50 | SMEAtten: Fast and Memory-Efficient Outer Product-based Attention on ARMv9 CPUs with SME (abstract) |
| 12:15 | HGNNMap: Heterogeneous Graph Neural Network-Based Mapping for Spatial Accelerators (abstract) PRESENTER: Shengzhong Tang |
| 11:00 | An Event-based Causality Model for Mobile Multiagent Systems (abstract) |
| 11:25 | Efficient Parallel Algorithms for Hypergraph Matching (abstract) |
| 11:50 | Minimizing Communication Costs in Inner Product Toom-Cook Algorithms (abstract) |
| 12:15 | Asynchronous Checkpoint for Eventually Consistent Databases (abstract) |
| 14:00 | Abstraction at What Cost? Evaluating Python DSLs for Transformer GPU Kernel Primitives on Nvidia Blackwell (abstract) |
| 14:25 | Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors (abstract) |
| 14:50 | Performance Portable BLAS3 Micro-kernel Generator (abstract) |
| 15:15 | How Efficient Are the Efficient Cores? - An Experimental Evaluation of Energy Efficiency in Asymmetric Multicore Processors (abstract) |
| 14:00 | Static Scheduling of DAG Workflows on Homogeneous Parallel Platforms Using Longest Betweenness Centrality (abstract) |
| 14:25 | ApproxRAG: A Systematic Framework for Mitigating I/O Overheads in Retrieval-Augmented Generation (abstract) |
| 14:50 | Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters (abstract) |
| 15:15 | Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters (abstract) PRESENTER: Wei Da |
| 14:00 | Lessons Learned on Program Recognition Robustness under Code Transformations (abstract) PRESENTER: Markus Puura |
| 14:10 | Reliability of Programming Languages: Sustainability inside the Codebase (abstract) PRESENTER: Darius Neațu |
| 14:20 | Changing the Future - Temporal Pointers as Unifying Dependency Abstraction (abstract) |
| 14:30 | The Invisible GPU: Transparent Stencil Acceleration from Cellular Automata to Fortran Atmospheric Models (abstract) |
| 14:40 | Epileptic Seizure Detection on Ultra-Low Power RISC-V SoC (abstract) PRESENTER: David Díaz-Reyes |
| 14:50 | Task-Parallel, Mixed-Precision Hybridized Solvers for Earthquake-Cycle Simulation (abstract) |
| 15:00 | Improving the Efficiency of Data Processing and Analysis in High-Energy Physics (abstract) PRESENTER: Florine Willemijn de Geus |
| 15:10 | Learning Performance and Energy Proxy from Hardware Performance Counters for Low-Overhead Autotuning (abstract) PRESENTER: Reilta Christine Dantas Maia |
| 15:20 | Beyond Prediction Accuracy: Understanding Runtime Estimation in HPC Scheduling (abstract) PRESENTER: Davide Leone |
| 15:30 | Pareto-based Workflow Scheduling for Makespan and Data Transfer Volume Co-optimisation (abstract) PRESENTER: Yani Ping |
| 14:00 | RWIP: Region-Level Write-Intensity Prediction for GPU L2 Cache based on Hybrid-Retention STT-MRAM (abstract) |
| 14:25 | Characterizing GPU Warm-Up Beyond Data Transfers (abstract) |
| 14:50 | Taking the Leap: Efficient and Reliable Fine-Grained NUMA Migration in User-space (abstract) |
| 15:15 | Performance Analysis of Hardware-Accelerated Compressed Memory Swap (abstract) |
| 16:30 | CoDA: Constraint-Based Distance Adjustment Optimization for Distance-Based ISAs (abstract) |
| 16:55 | Assessing the Performance Impact of Data Layouts: a Benchmarking Approach (abstract) |
| 17:20 | A Compiler-Assisted Workflow for Efficiency-Guided Selective Tracing (abstract) |
| 16:30 | MARS: Multi-model Aware Real-time Scheduler for NPU-coordinated DLI Tasks (abstract) |
| 16:55 | S-CQR: Stratified Calibration for Runtime Prediction in HPC Backfill Scheduling (abstract) |
| 17:20 | Coordinated Resource Management for Energy-Efficient DNN Inference on Heterogeneous Edge Devices (abstract) |
| 16:30 | RAGNN: A Resource-Aware System for Graph Neural Network Training at Scale (abstract) |
| 16:55 | Two-Stage Hierarchy-Aware Learning with Gradient Conflict Mitigation for HLS Latency and Resource Prediction (abstract) |
| 17:20 | DACOS: Dependency-Aware Cross-Kernel Overlapping for Optimizing Short-Sequence Workloads in LLM Applications (abstract) PRESENTER: Zhaoyang Hao |
| 16:30 | Performance evaluation of a CPU-GPU coprocessing-based simulation software on converged computing architectures (abstract) |
| 16:55 | How much slack is in a multiprocessor schedule? (abstract) |
| 17:20 | Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication (abstract) PRESENTER: Huimin Liu |
View this program: with abstractssession overviewtalk overview
Decentralized Learning at the Edge: Unlocking the Future of AI
As concerns around data privacy grow, decentralized learning enables collaborative model training without sharing raw data. Yet, exchanging model updates can still leak sensitive information, and system performance is often hindered by stragglers in heterogeneous environments.
In this talk, I will highlight these key challenges and present recent advances in straggler-resilient learning and privacy-preserving protocols. Together, these results pave the way for decentralized learning systems that are both efficient and secure, enabling scalable AI at the edge.
| 10:30 | Constraint Driven Global Data Layout Optimization for Tensor Expressions (abstract) |
| 10:55 | Understanding Power Limiting Mechanisms in Modern Processors: A Deep Dive into Intel RAPL and Turbo Boost Dynamics (abstract) PRESENTER: Romial Menra |
| 11:20 | CASTM: An API for Accelerating Zero-Knowledge Proof Kernels on CGRA Architectures (abstract) |
| 11:45 | Accelerating Sharded Data Parallelism at Scale with Federated Learning (abstract) PRESENTER: Gianluca Mittone |
| 12:10 | Engineering Scalable Distributed List Ranking (abstract) |
| 12:35 | Accelerating Multi-Agent Reinforcement Learning on Heterogeneous Platforms (abstract) |
| 14:00 | Repurposing the Cross-Segment Space on Sunway SW26010Pro for MC-Balanced Bigshare Execution (abstract) PRESENTER: Qixin Chang |
| 14:25 | NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures (abstract) |
| 14:50 | SISA: A Scale-In Systolic Array for GEMM Acceleration (abstract) |
| 15:15 | Realizable N:M Sparse Transformer Inference via Search–Kernel Co-Design (abstract) PRESENTER: Yiming Liu |
| 14:00 | Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs (abstract) |
| 14:25 | When Faster Kernels Do Not Mean Faster Applications in Exascale GPU Systems (abstract) |
| 14:50 | Optimizing SpMV Kernel for Banded Matrices on FPGAs Using High-Level Synthesis (abstract) |
| 15:15 | OmniZK: A Versatile Accelerator Architecture for Zero-Knowledge Proofs (abstract) |
| 14:00 | Data Layout Transformation for Memory Performance Optimisation (abstract) PRESENTER: Jolly Chen |
| 14:10 | From Functions to Complex Applications: Toward High-Performance Serverless Computing (abstract) PRESENTER: Valerio Besozzi |
| 14:20 | Toward Adaptable Serverless Workflows: A Runtime-Adjustable DAG Approach for Streaming Pipelines (abstract) PRESENTER: Matteo Della Bartola |
| 14:30 | Collaborative Agentic AI in the Internet-of-Agents (abstract) PRESENTER: Saul Urso |
| 14:40 | Toward Efficient Community Detection on Time-Varying Attributed Graphs (abstract) PRESENTER: Nelson Aloysio Reis de Almeida Passos |
| 14:50 | AI-driven Data Management in Distributed IoT–Edge–Cloud Systems (abstract) PRESENTER: Adrianna Bodziony |
| 15:00 | Decentralized Federated Learning for Net Active Power Probabilistic Forecasting in Smart Microgrids (abstract) PRESENTER: Mohsen Salimi Khanghah |
| 15:10 | Towards a General Analytical Theory of LLM Inference Energy Efficiency (abstract) PRESENTER: Andrea Proia |
| 15:20 | Making Privacy-Preserving Computation Practical for Data Center Coordination and Language Model Inference (abstract) PRESENTER: Seyda Nur Guzelhan |
| 14:00 | A Brief Survey of Task-Level Resilience Techniques (abstract) |
| 14:20 | Bridging Classical High‑Performance and Quantum Computing: Lessons, Challenges, and Opportunities (abstract) |
| 14:40 | Intelligent resource management in distributed infrastructures (abstract) |
| 15:00 | An Intermediate Representation for Heterogeneous Environments (abstract) |
| 15:20 | Lost in Translation? Experiences Building an LLM for Compiler IR Interoperability (abstract) |
| 16:30 | Lyapunov-Guided KV Cache Control for Multi-Tenant Edge LLM Servicing (abstract) PRESENTER: Haishuo Yu |
| 16:55 | FlowGPU: Transparent and Efficient GPU Checkpointing and Restore (abstract) |
| 17:20 | Adaptive Fault Tolerance in Time-Sensitive Networking (abstract) |
| 16:30 | AtSpMV: Model-Guided Adaptive Tiling and Load Balancing for SpMV on GPUs (abstract) PRESENTER: Cheng Jiajun |
| 16:55 | A Task Parallel Algorithm For Fast Hybridized PDE Solvers (abstract) |
| 17:20 | GWBP: Accelerating Weighted Back-Projection for Image Reconstruction via Efficient GPU Parallel Optimization (abstract) |
| 16:30 | Constella: A Novel Framework for Cost-Efficient Distributed AI Inference in LEO Space Data Centers (abstract) |
| 16:55 | Evaluating the Parallelization Capabilities of State-of-the-art Agentic Large Language Models (abstract) |
| 17:20 | CCGS: A Cross-Modal Collaborative Gradient Sparsification for Accelerating Distributed Multimodal Model Training (abstract) |
| 16:30 | Co-Designing AI & HPC Supercomputers as Grid-Responsive, Carbon-Aware Citizens (abstract) |
| 16:50 | Beyond efficiency: feedback-based resource management to limit computing impacts (abstract) |
| 17:10 | RISC-V for HPC: How Ready Are We? (abstract) |
View this program: with abstractssession overviewtalk overview
Sustainable, Pervasive Agentic AI - Nowhere to go, but UP!
Computing is now dominated by generative AI, with focus rapidly shifting from training to (agentic) inference. The fast pace imposed by generative AI scaling laws requires sustained energy efficiency improvements, which cannot be matched simply by classical (Moore’s Law) technology evolution. We must move toward full three-dimensional scaling: UP, in stacking and in scale, is the only way to go. To tackle the challenge, we need to aggressively optimize computing platforms leveraging specialization. However, the design of scalable domain-specific 3D architectures requires actionable understanding of Amdahl’s, Patterson’s, and Little’s laws. In this talk, I will show how to leverage these fundamental laws in computer engineering to design the next generation of AI chips and systems.
10:30 E4 Computer Engineering — Roberto Rocco
10:40 GreenShift — Horacio González-Vélez
10:50 DDN — Thomas Blum
11:00 ICSC — Ugo Becciani
11:10 EUPEX — Pélagie Alves
11:20 Bull — Pélagie Alves
11:30 Huawei — Chong Li and Hongxing Wang
| 11:45 | SparX: Cache-Aware Sparse Matrix Multiplication Kernels on AMX-based CPUs (abstract) |
| 12:10 | Fast Checkpointing in Disaggregated Persistent Memory (abstract) |
| 12:35 | Efficient Parallel Enumeration of Disjoint Minimal Failure-Inducing Subsets (abstract) |
| 11:45 | G-STAR: GPU-Accelerated Statistical Static Timing Analysis using Level-by-level Replication (abstract) |
| 12:10 | JAX vs Kokkos Performance Comparison: A Magnetohydrodynamic Solver Case Study (abstract) |
| 12:35 | A Novel Parallel Strategy for Monte Carlo Simulations in Heterogeneous Distributed Systems (abstract) |
| 11:45 | Harmonia: QoS-Aware and High-Throughput Generative Inference with a Single GPU (abstract) |
| 12:10 | DAG-P: Efficient Fine-Grained Partitioning of Open-Source Transformer Models via Dependency-Aligned Planning (abstract) |
| 12:35 | Layer-wise CPU-GPU Scheduling for LLM Inference on Memory-Limited Consumer GPUs (abstract) |
| 11:45 | Throughput Optimization for Multi-level Speculative Decoding (abstract) |
| 12:10 | SuperSFL: Resource-Heterogeneous Federated Split Learning with Weight-Sharing Supernet (abstract) PRESENTER: Abdullah Al Asif |
| 12:35 | Adaptive Data Augmentation with Bayesian Optimization for Basic Block Throughput Prediction (abstract) |
| 14:00 | Moirai: Dependency-Impact-Based Communication Scheduling for Multi-Job Distributed Deep Learning Cluster (abstract) PRESENTER: Jingbin Yang |
| 14:25 | SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph Pipelines (abstract) |
| 14:50 | PeakLife: Proactive VM Management via Joint Forecasting (abstract) PRESENTER: Bowen Sun |
| 14:00 | Thunder: Efficient Multi-Node FHE Acceleration Framework via In-Transit Computation (abstract) |
| 14:25 | X-BQSR: Holistic Acceleration of Base Quality Recalibration for Scalable Genomic Analysis (abstract) |
| 14:50 | Privacy-Preserving Data Center Demand Response using Multi-Party Computation (abstract) PRESENTER: Seyda Nur Guzelhan |
| 14:00 | TAIN: Scalable Tensor-Aware Acceleration for Quantum Chemistry on Heterogeneous Architectures (abstract) PRESENTER: Wenhao Liang |
| 14:25 | Hierarchical Parallel Computing for Optimal Qubit Mapping in NISQ Systems (abstract) |
| 14:50 | QSplit: A Workflow-Oriented Hybrid Quantum-Classical Optimization Framework (abstract) |
| 14:00 | MatrixFold: Unleashing Manycore CPUs with Outer-Product Units for Mixed-Precision AlphaFold Inference (abstract) |
| 14:25 | PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving (abstract) |
| 14:50 | Performance Characterization and Optimization of LLM Inference on Tenstorrent AI Accelerators (abstract) |