SBAC-PAD 2026: 38TH IEEE/SBC INTERNATIONAL SYMPOSIUM ON COMPUTER ARCHITECTURE AND HIGH PERFORMANCE COMPUTING (SBAC-PAD)
PROGRAM

Days: Wednesday, October 14th Thursday, October 15th Friday, October 16th

Wednesday, October 14th

View this program: with abstractssession overviewtalk overview

09:30-10:30 Session 2: Keynote: Estela Suarez

Strategies for energy efficiency in HPC
Energy efficiency needs to be addressed in HPC at all layers of the stack, from application development up to hardware deployment and system operation. While significant gains have been achieved through advances in hardware architecture, improving how existing systems are operated remains an equally important opportunity. This presentation offers a broad overview of established approaches to improving energy efficiency, including hardware innovations, system software techniques, and application settings. It will then focus on operational practices aiming at obtaining the maximum throughput out of existing infrastructures, at the minimum possible energy cost. In particular, it discusses how runtime decisions based on system state and workload characteristics can be leveraged to reduce energy waste while maintaining performance. The presentation draws on work carried out in the SEANERGYS project, which develops a software suite combining monitoring, data analytics, and scheduling to achieve better energy-aware system operations.

Location: Lecture Hall
10:30-11:30Coffee Break
11:00-13:00 Session 3A: Computer Architecture (I)
Location: Lecture Hall
11:00
Chanaka Hettige (Indiana University Bloomington, United States)
Martin Swany (Indiana University Bloomington, United States)
Offloading Floating-Point Based Collective Operations to Programmable Logic for Distributed Learning (abstract)
PRESENTER: Martin Swany
11:30
Víctor Galindo-Garre (Universidad de Murcia, Spain)
Rubén Titos-Gil (Universidad de Murcia, Spain)
Ricardo Fernández-Pascual (Universidad de Murcia, Spain)
Alberto Ros (Universidad de Murcia, Spain)
Analysis and Implementation Alternatives of Multi-grain Coherence Directories (abstract)
12:00
Munawira Kotyad (Indian Institute of Technology Bombay; Pillai University, India)
Ayush Agrawal (Indian Institute of Technology, Bombay, India)
Virendra Singh (Indian Institute of Technology, Bombay, India)
HOOP: Hint based Out of Order Processing for GPGPUs via Per-Kernel Dependency Analysis (abstract)
PRESENTER: Munawira Kotyad
12:30
Marwan Fetteha (McMaster University, Canada)
Ameer Abdelhadi (McMaster University, Canada)
Hardware Design of a Low-Cost, DSP-Free, Spiking Neural Network for Edge Devices (abstract)
11:00-13:00 Session 3B: Performance Evaluation (I)
11:00
Marc Jordà (Barcelona Supercomputing Center (BSC), Spain)
Antonio J. Peña (Barcelona Supercomputing Center (BSC), Spain)
Assessing Arm SPE for Automatic Data Placement in Heterogeneous Memory Systems (abstract)
11:30
André Sacilotto Santos (PUCRS, Brazil)
Miguel Gomes Xavier (PUCRS, Brazil)
Sören Becker (TU Berlin, Germany)
Cesar Augusto Fonticielha De Rose (PUCRS, Brazil)
Odej Kao (TU Berlin, Germany)
Cross-Application Interference Profiling in Consolidated Cloud Servers: Improving Portability, Overhead and Fidelity in IntP (abstract)
12:00
Ana Clara Lannes (University of São Paul, Brazil)
Bernardo Pereira (University of São Paulo, Brazil)
Alfredo Goldman (University of São Paulo, Brazil)
Optimization of Resource Efficiency in HPC: Exploiting Reduced Precision in Geophysical Simulations (abstract)
11:00-13:00 Session 3C: WCC
11:00
Garvit Mittal (Bharat Electronics Limited, India)
Sahil Tomar (Bharat Electronics Limited, India)
Sandeep Kumar (Bharat Electronics Limited, India)
QYOLO: Lightweight Object Detection via Quantum Inspired Shared Channel Mixing (abstract)
11:30
Zeba Mahmood (University of Applied Sciences and Arts Dortmund, Germany)
Stephan Recker (University of Applied Sciences and Arts Dortmund, Germany)
The Volatility Tax: Quantifying Token-Escrow Risk in Decentralised Compute Markets (abstract)
12:00
Fernando Guzman (Pontifical Catholic University of Peru, Peru)
Cesar Santivanez (Pontifical Catholic University of Peru, Peru)
Calibrated Multi-Step CPU Usage Forecasting via Fractional Brownian Motion-Driven FARIMA for Proactive VM Migration (abstract)
13:00-14:00Lunch Break
14:00-15:00 Session 4: Keynote: Viktor K. Prasanna

Towards Real-time AI at the Edge

Edge AI applications including Agentic AI systems require real-time performance. In addition, in many applications, energy efficiency becomes a critical metric. To meet these requirements, recently many edge platforms have been proposed including GPUs, FPGAs, NPUs and heterogeneous architectures such as AI PCs. These devices are being used along with multi-core and novel memory technologies to realize advanced platforms to accelerate edge inference and support diverse applications requiring low latency. We will review emerging technologies for edge inference and advances in reconfigurable computing over the past three decades leading up to current innovations in FPGA accelerators for AI. We will illustrate parallel architectures and algorithms for commercial, defense and space applications. Using our algorithm-architecture co-design methodology to realize high performance accelerators for these applications, we demonstrate the role of modeling and algorithmic optimizations to develop highly efficient Intellectual Property (IP) cores for FPGAs and realize end to end application acceleration. We illustrate our methodology by developing high performance designs for graph machine learning, long context LLM inference, Mixture of Agents inference and SAR ATR. We conclude by identifying opportunities and challenges in exploiting emerging heterogeneous architectures composed of multi-core processors, FPGAs, integrated GPUs, NPUs and accelerators.

Location: Lecture Hall
15:00-15:30Coffee Break
15:30-17:30 Session 5A: Computer Architecture (II)
Location: Lecture Hall
15:30
Tamara Lugo (Universidad Carlos III de Madrid, Spain)
Javier Fernandez (University Carlos III of Madrid, Spain)
Jesús Carretero (University Carlos III of Madrid, Madrid, Spain)
New techniques to reduce cache interference in real-time mixed-criticality systems on multicore platforms (abstract)
16:00
Yusuf Tekin (Istanbul Technical University, Turkey)
Berna Örs (Istanbul Technical University, Turkey)
Implementation and Acceleration of MLP-Based IDS for FPGA (abstract)
16:30
Philippos Papaphilippou (University of Southampton, UK)
LUTstructions: Fast-Reconfigurable FPGA-Based Instructions (abstract)
17:00
Marco Ho (School of Computing & Academic Studies, British Columbia Institute of Technology, Canada)
Michael Hsiao (Virginia Tech, United States)
Jeeho Ryoo (Fairleigh Dickinson University, Canada)
ARTA: Adaptive Reinforcement-Learning-Based Throttling Agent for RowHammer Vulnerabilities (abstract)
15:30-17:30 Session 5B: Performance Evaluation (II)
15:30
Amritanshu Verma (Johannes Gutenberg University, Mainz, Germany)
Sarah Neuwirth (Johannes Gutenberg University, Mainz, Germany)
Beyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on modern GPUs (abstract)
16:00
Matheus Costa (Federal University of Rio Grande do Sul, Brazil)
Sandro Rigo (UNICAMP, Brazil)
Antigoni Georgiadou (Oak Ridge National Laboratory, United States)
Bronson Messer (Oak Ridge National Laboratory, United States)
Philippe Navaux (Federal University of Rio Grande do Sul, Brazil)
Silvio Rizzi (Argonne National Laboratory, United States)
Arthur Lorenzon (Federal University of Rio Grande do Sul, Brazil)
Predicting the Effectiveness of GPU Sharing for AI Inference on Aurora (abstract)
16:30
Francesco Sgherzi (SiPearl, Spain)
Nicolas Bouton (SiPearl, France)
Julie Gaspar (SiPearl, France)
Clement Gavoille (SiPearl, France)
Alexis Laplanche (SiPearl, France)
Vijendra Singh (SiPearl, Spain)
Etienne Renault (SiPearl, France)
YABLT: Modeling Memory Bandwidth Effects by Limiting Bandwidth Availability (abstract)
15:30-17:30 Session 5C: WAMCA
15:30
Minh Chau Nguyen (University of Toronto Scarborough, Canada)
Marcelo Ponce (University of Toronto Scarborough, Canada)
A Comparative Study of Parallel Implementations of Short-Time Fourier Transform (abstract)
16:00
Daniel Benedict (Texas Tech University, United States)
Between Customization and Constraint: Performance, Scalability, and Software Readiness of RISC-V for HPC (abstract)
16:30
Sairo Santos (UFERSA - Universidade Federal Rural do Semi-árido, Brazil)
Rodrigo Machniewicz Sokulski (UFPR - Universidade Federal do Paraná, Brazil)
Tiago Rodrigo Kepe (Federal institute of Paraná, Brazil)
Marco Antonio Zanata Alves (UFPR - Universidade Federal do Paraná, Brazil)
Vector-In-Memory Architecture for Data-centric Applications (abstract)
17:00
Alberto Cascajo (University Carlos III of Madrid, Spain)
Javier Fernandez Muñoz (University Carlos III of Madrid, Spain)
David E. Singh (University Carlos III of Madrid, Madrid, Spain)
Jesús Carretero (University Carlos III of Madrid, Madrid, Spain)
A Semi-Autonomous Monitoring Environment for Malleable HPC Systems (abstract)
20:00-22:00Welcome cocktail

Location: U-Music hotel, C. de la Paz, 11

UMusic Hotel Madrid represents the transformation of historical heritage into a pioneering concept of musical hospitality, bringing together the iconic Albéniz Theater and the former Hotel Madrid in a single space.

The Albéniz Theater, originally opened in 1945 as a landmark of Madrid's lyrical and theatrical scene, remained closed for years following threats of demolition. Declared a Property of Cultural Interest in 2016 thanks to civic and cultural mobilization, the venue was renovated through an investment of nearly 30 million euros driven by SOCIMI Silicius and operated by UMusic Hotels (a division of Universal Music Group), reopening its doors in late 2022.

The complex does not function as a simple accommodation with themed decor, but as a space where gastronomy, five-star lodging, and performing arts converge.

Thursday, October 15th

View this program: with abstractssession overviewtalk overview

09:00-10:00 Session 6: Distributed Systems, Networking and Storage
Location: Lecture Hall
09:00
Saeideh Alinezhad (imec, Belgium)
Ramzi Baaguigui (imec, Belgium)
Kalliopi Tzimi (imec, Belgium)
Tommaso Marinelli (imec, Belgium)
Konstantinos Tovletoglou (imec, Belgium)
Dissecting End‑to‑End SSD I/O Latency in Storage Systems: Cross‑Layer Bottlenecks and Technology Trade‑offs (abstract)
09:30
Théo Jolivel (Inria, France)
François Tessier (Inria, France)
Jakob Luettgau (Inria, France)
Gabriel Antoniu (Inria, France)
Philippe Deniel (CEA, France)
A Methodology for System-Scale I/O Pattern Taxonomy for HPC Workloads (abstract)
PRESENTER: Théo Jolivel
10:00-10:30Coffee Break
10:30-11:30 Session 7: Keynote: Domenico Talia

Distributed Machine Learning: Bridging Cloud and Edge Systems
Edge-cloud solutions are being used to collect and analyze large amounts of data generated by IoT devices in various application domains, such as urban mobility, smart cities, healthcare, and augmented reality. We must be able to combine techniques and algorithms of data analysis and machine learning with the scalable architectures of Cloud systems and Edge technologies. This approach can reduce latency and network congestion associated with traditional cloud-based machine learning techniques by processing data locally on edge devices before sending it to the cloud for further analysis. This keynote discusses distributed machine learning and proposes a reference architecture to adapt distributed machine learning algorithms at the edge-cloud continuum. Real applications are presented, and the main open research issues are discussed.

Location: Lecture Hall
12:30-13:30Lunch Break
13:30-15:30 Session 9: Best Paper Session
Location: Lecture Hall
13:30
Luc Joffily Ribas (Universidade Estadual de Campinas, Brazil)
Guido Costa Souza de Araújo (Universidade Estadual de Campinas, Brazil)
Cross-Platform MLIR Based GEMM Micro-Kernel Generation (abstract)
14:00
Rappy Saha (University of Glasgow, UK)
Nima Amirafshar (Heidelberg University, Germany)
Jude Haris (University of Glasgow, UK)
Nima Taherinejad (Heidelberg University, Germany)
José Cano (University of Glasgow, UK)
FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining (abstract)
14:30
Caio Salvador Rohwedder (AMD, Canada)
João Paulo Labegalini de Carvalho (AMD, Canada)
Jose Nelson Amaral (University of Alberta, Canada)
Zero-Copy GEMM-Based Fast Convolution (abstract)
15:00
Resit Sendag (University of Rhode Island, United States)
Matthew Constant (University of Rhode Island, United States)
The Fallacy of Independent Ceilings: Characterizing Coupled Load-Branch Stall Interaction (abstract)
15:30-16:00Coffee Break
16:00-17:30 Session 10: Parallel Applications and Algorithms
Location: Lecture Hall
16:00
Vinícius Daniel Spadotto (UFRGS, Brazil)
Lucas Mello Schnorr (UFRGS, Brazil)
MPI-Based 3D Anisotropic RTM with Fletcher's Method and Ghost Cell Halo Exchange (abstract)
16:30
Tom Budon (LIFO - University of Orléans, France)
Florent De Martin (BRGM, France)
Sylvain Jubertie (LIFO - University of Orléans, France)
Sébastien Limet (LIFO - University of Orléans, France)
Emmanuel Melin (LIFO - University of Orléans, France)
Shuaitao Wang (BRGM, France)
Speedup and Energy benchmarks of GPU and SIMD spectral finite-element method (abstract)
17:00
Fukuharu Tanaka (Osaka University, Japan)
Hamada Rizk (The University of Osaka, RIKEN Center for Computational Science, Japan)
Moustafa Youssef (The American University in Cairo, Egypt)
Hirozumi Yamaguchi (The University of Osaka, RIKEN Center for Computational Science, Japan)
Resource-Aware Model Selection for Scalable Indoor Localization on HPC Platforms (abstract)
17:30
Hector Migallon (University Miguel Hernández, Spain)
Miguel Soria Zaragoza (University Miguel Hernández, Spain)
Adrián Sarrías (University Miguel Hernández, Spain)
Miguel Martínez Rach (University Miguel Hernández, Spain)
Otoniel López Granado (University Miguel Hernández, Spain)
Architecture-Aware GPU Acceleration for Scalable Real-Time Industrial Hyperspectral Classification (abstract)
PRESENTER: Adrián Sarrías
20:00-22:00Banquet

Location: Casa Suecia, Marqués de Casa Riera, 4

Inaugurated in 1956 next to the Círculo de Bellas Artes, Casa Suecia was born as the cultural, diplomatic, and social epicenter of the Scandinavian community in Madrid. Designed by architect Mariano Garrigues, the building originally housed the Scandinavian Center and an exclusive hotel.
The project emerged with the aim of strengthening commercial and cultural ties between Spain and Nordic countries. It was patronized in its early years by King Gustaf VI Adolf of Sweden, immediately becoming a vital meeting point for diplomacy, business, and the Swedish community in the capital.
Following a comprehensive renovation in the early 21st century, the building reopened incorporating the NH Collection Madrid Suecia hotel. The dining and leisure space retained the name Casa Suecia, preserving its historical legacy across several areas:

  • Cocktail Bar Hemingway: A speakeasy-style cocktail bar in the basement—hidden behind a door in the restroom area—with 1950s-inspired decor.
  • La Terraza: A popular rooftop terrace offering 360-degree views over the city center rooftops.
  • The Restaurant: A Mediterranean cuisine space that maintains nods to traditional Swedish gastronomy.
Friday, October 16th

View this program: with abstractssession overviewtalk overview

09:00-10:30 Session 11A: System and Software
Location: Lecture Hall
09:00
Jie Lei (Universitat Politècnica de València, Spain)
Héctor Martínez (Universidad de Córdoba, Spain)
Adrián Castelló (Universitat Politècnica de València, Spain)
Automatic Generation of Portable RVV Micro-Kernels through a Hybrid MLIR--xDSL Compilation Framework (abstract)
PRESENTER: Jie Lei
09:30
Paula Sánchez-Checa (Universidad Carlos III de Madrid, Spain)
Jesus Carretero (Universidad Carlos III de Madrid, Spain)
David E. Singh (Universidad Carlos III de Madrid, Spain)
Leveraging Dynamic Resource Management for Energy Efficiency: To Speed-up or To Green-up (abstract)
10:00
Hamed Sharifian (Queen's University, Canada)
Amirhossein Sojoodi (Queen's University, Canada)
Ahmad Afsahi (Queen's University, Canada)
Exploiting Virtual Topology for Load- and Communication-Aware Neighborhood Collectives (abstract)
10:30
Fabio Souza (INESC TEC and University of Minho, Portugal)
Daniel Sodré (Universidade Federal Fluminense, Brazil)
João Paulo (INESC TEC and University of Minho, Portugal)
Ricardo Macedo (INESC TEC and University of Minho, Portugal)
Cristina Boeres (Universidade Federal Fluminense, Brazil)
Vinod Rebello (Universidade Federal Fluminense, Brazil)
Felipe Portella (Petróleo Brasileiro S.A. (Petrobras), Brazil)
Paulo Estrela (Petróleo Brasileiro S.A. (Petrobras), Brazil)
Renzo Malini (Petróleo Brasileiro S.A. (Petrobras), Brazil)
Phoning Home: Repatriating output from cloud jobs with a user-level, provider-agnostic data transfer tool (abstract)
10:30-11:00Coffee Break
11:00-12:00 Session 12: Keynote: Rosa M. Badia

Tasks Based Hybrid Workflows for Emergent Technologies

The Barcelona Supercomputing Center hosts a twin installation based on transmon qubits, superconducting technology funded by Spain. Another annealing system co-funded by EuroHPC JU and Spain is also expected. Significant research and development is done by BSC researchers from the Quantic groups in topics related to quantum computing, which are closer to its physics nature. Complementary to this research, and in collaboration with our colleagues, the Workflows and Distributed Computing group is doing research activities related to software aspects related to the integration of HPC + Quantum Computing (QC).

The talk will describe the steps taken towards integration of HPC and QC at the BSC. We will describe how a traditional workflow environment is being extended to support hybrid HPC+quantum computing workflows.The current solution enables the hybrid execution of workflows in CPUS, GPUs and Quantum systems (onsite and through cloud access). The talk will present these topics and our plans for the near future.

Location: Lecture Hall
12:00-13:30 Session 13A: Parallel and Distributed Computing for AI and Data Analytics
Location: Lecture Hall
12:00
Máximo Rodríguez Herrero (Universidad Carlos III, Spain)
Dante D. Sanchez-Gallegos (Universidad Carlos III de Madrid, Spain)
Jesus Carretero (Universidad Carlos III de Madrid, Spain)
GPU-Efficient XAI Uncertainty Estimation for nnUNet (abstract)
12:30
Ferran Soler (Universitat Jaume I, Spain)
Jose I. Mestre (Universitat Politècnica de València, Spain)
Manuel F. Dolz (Universitat Jaume I, Spain)
José I. Aliaga (Universitat Jaume I, Spain)
Design and Evaluation of Continuous Synthetic Data Pipelines for Vision Pre-training (abstract)
PRESENTER: Ferran Soler
13:00
Hassan Mohsen (HPC, France)
Parallelizing Deep Learning: A Unified Study of CPU and GPU Partitioning Strategies, Bottlenecks, and Design Principles (abstract)
13:30
Hector Romero Ugalde (Bull, France)
Youssef Faqir-Rhazoui (BULL, Spain)
Marlon Funk (BULL, Spain)
Jesus Gorroñogoitia Cruz (Bull, Spain)
Evaluating Runtime Prediction and Conservative Walltime Policies in HPC Systems (abstract)