Days: Wednesday, October 14th Thursday, October 15th Friday, October 16th
View this program: with abstractssession overviewtalk overview
Strategies for energy efficiency in HPC
Energy efficiency needs to be addressed in HPC at all layers of the stack, from application development up to hardware deployment and system operation. While significant gains have been achieved through advances in hardware architecture, improving how existing systems are operated remains an equally important opportunity. This presentation offers a broad overview of established approaches to improving energy efficiency, including hardware innovations, system software techniques, and application settings. It will then focus on operational practices aiming at obtaining the maximum throughput out of existing infrastructures, at the minimum possible energy cost. In particular, it discusses how runtime decisions based on system state and workload characteristics can be leveraged to reduce energy waste while maintaining performance. The presentation draws on work carried out in the SEANERGYS project, which develops a software suite combining monitoring, data analytics, and scheduling to achieve better energy-aware system operations.
| 11:00 | Offloading Floating-Point Based Collective Operations to Programmable Logic for Distributed Learning (abstract) PRESENTER: Martin Swany |
| 11:30 | Analysis and Implementation Alternatives of Multi-grain Coherence Directories (abstract) |
| 12:00 | HOOP: Hint based Out of Order Processing for GPGPUs via Per-Kernel Dependency Analysis (abstract) PRESENTER: Munawira Kotyad |
| 12:30 | Hardware Design of a Low-Cost, DSP-Free, Spiking Neural Network for Edge Devices (abstract) |
| 11:00 | Assessing Arm SPE for Automatic Data Placement in Heterogeneous Memory Systems (abstract) |
| 11:30 | Cross-Application Interference Profiling in Consolidated Cloud Servers: Improving Portability, Overhead and Fidelity in IntP (abstract) PRESENTER: André Sacilotto Santos |
| 12:00 | Optimization of Resource Efficiency in HPC: Exploiting Reduced Precision in Geophysical Simulations (abstract) |
| 11:00 | QYOLO: Lightweight Object Detection via Quantum Inspired Shared Channel Mixing (abstract) |
| 11:30 | The Volatility Tax: Quantifying Token-Escrow Risk in Decentralised Compute Markets (abstract) |
| 12:00 | Calibrated Multi-Step CPU Usage Forecasting via Fractional Brownian Motion-Driven FARIMA for Proactive VM Migration (abstract) |
Towards Real-time AI at the Edge
Edge AI applications including Agentic AI systems require real-time performance. In addition, in many applications, energy efficiency becomes a critical metric. To meet these requirements, recently many edge platforms have been proposed including GPUs, FPGAs, NPUs and heterogeneous architectures such as AI PCs. These devices are being used along with multi-core and novel memory technologies to realize advanced platforms to accelerate edge inference and support diverse applications requiring low latency. We will review emerging technologies for edge inference and advances in reconfigurable computing over the past three decades leading up to current innovations in FPGA accelerators for AI. We will illustrate parallel architectures and algorithms for commercial, defense and space applications. Using our algorithm-architecture co-design methodology to realize high performance accelerators for these applications, we demonstrate the role of modeling and algorithmic optimizations to develop highly efficient Intellectual Property (IP) cores for FPGAs and realize end to end application acceleration. We illustrate our methodology by developing high performance designs for graph machine learning, long context LLM inference, Mixture of Agents inference and SAR ATR. We conclude by identifying opportunities and challenges in exploiting emerging heterogeneous architectures composed of multi-core processors, FPGAs, integrated GPUs, NPUs and accelerators.
| 15:30 | New techniques to reduce cache interference in real-time mixed-criticality systems on multicore platforms (abstract) |
| 16:00 | Implementation and Acceleration of MLP-Based IDS for FPGA (abstract) |
| 16:30 | LUTstructions: Fast-Reconfigurable FPGA-Based Instructions (abstract) |
| 17:00 | ARTA: Adaptive Reinforcement-Learning-Based Throttling Agent for RowHammer Vulnerabilities (abstract) |
| 15:30 | Beyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on modern GPUs (abstract) |
| 16:00 | Predicting the Effectiveness of GPU Sharing for AI Inference on Aurora (abstract) |
| 16:30 | YABLT: Modeling Memory Bandwidth Effects by Limiting Bandwidth Availability (abstract) |
| 15:30 | A Comparative Study of Parallel Implementations of Short-Time Fourier Transform (abstract) |
| 16:00 | Between Customization and Constraint: Performance, Scalability, and Software Readiness of RISC-V for HPC (abstract) |
| 16:30 | Vector-In-Memory Architecture for Data-centric Applications (abstract) |
| 17:00 | A Semi-Autonomous Monitoring Environment for Malleable HPC Systems (abstract) |
Location: U-Music hotel, C. de la Paz, 11
UMusic Hotel Madrid represents the transformation of historical heritage into a pioneering concept of musical hospitality, bringing together the iconic Albéniz Theater and the former Hotel Madrid in a single space.
The Albéniz Theater, originally opened in 1945 as a landmark of Madrid's lyrical and theatrical scene, remained closed for years following threats of demolition. Declared a Property of Cultural Interest in 2016 thanks to civic and cultural mobilization, the venue was renovated through an investment of nearly 30 million euros driven by SOCIMI Silicius and operated by UMusic Hotels (a division of Universal Music Group), reopening its doors in late 2022.
The complex does not function as a simple accommodation with themed decor, but as a space where gastronomy, five-star lodging, and performing arts converge.
View this program: with abstractssession overviewtalk overview
| 09:00 | Dissecting End‑to‑End SSD I/O Latency in Storage Systems: Cross‑Layer Bottlenecks and Technology Trade‑offs (abstract) |
| 09:30 | A Methodology for System-Scale I/O Pattern Taxonomy for HPC Workloads (abstract) PRESENTER: Théo Jolivel |
Distributed Machine Learning: Bridging Cloud and Edge Systems
Edge-cloud solutions are being used to collect and analyze large amounts of data generated by IoT devices in various application domains, such as urban mobility, smart cities, healthcare, and augmented reality. We must be able to combine techniques and algorithms of data analysis and machine learning with the scalable architectures of Cloud systems and Edge technologies. This approach can reduce latency and network congestion associated with traditional cloud-based machine learning techniques by processing data locally on edge devices before sending it to the cloud for further analysis. This keynote discusses distributed machine learning and proposes a reference architecture to adapt distributed machine learning algorithms at the edge-cloud continuum. Real applications are presented, and the main open research issues are discussed.
| 13:30 | Cross-Platform MLIR Based GEMM Micro-Kernel Generation (abstract) PRESENTER: Luc Joffily Ribas |
| 14:00 | FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining (abstract) |
| 14:30 | Zero-Copy GEMM-Based Fast Convolution (abstract) |
| 15:00 | The Fallacy of Independent Ceilings: Characterizing Coupled Load-Branch Stall Interaction (abstract) |
| 16:00 | MPI-Based 3D Anisotropic RTM with Fletcher's Method and Ghost Cell Halo Exchange (abstract) |
| 16:30 | Speedup and Energy benchmarks of GPU and SIMD spectral finite-element method (abstract) |
| 17:00 | Resource-Aware Model Selection for Scalable Indoor Localization on HPC Platforms (abstract) |
| 17:30 | Architecture-Aware GPU Acceleration for Scalable Real-Time Industrial Hyperspectral Classification (abstract) PRESENTER: Adrián Sarrías |
Location: Casa Suecia, Marqués de Casa Riera, 4
Inaugurated in 1956 next to the Círculo de Bellas Artes, Casa Suecia was born as the cultural, diplomatic, and social epicenter of the Scandinavian community in Madrid. Designed by architect Mariano Garrigues, the building originally housed the Scandinavian Center and an exclusive hotel.
The project emerged with the aim of strengthening commercial and cultural ties between Spain and Nordic countries. It was patronized in its early years by King Gustaf VI Adolf of Sweden, immediately becoming a vital meeting point for diplomacy, business, and the Swedish community in the capital.
Following a comprehensive renovation in the early 21st century, the building reopened incorporating the NH Collection Madrid Suecia hotel. The dining and leisure space retained the name Casa Suecia, preserving its historical legacy across several areas:
- Cocktail Bar Hemingway: A speakeasy-style cocktail bar in the basement—hidden behind a door in the restroom area—with 1950s-inspired decor.
- La Terraza: A popular rooftop terrace offering 360-degree views over the city center rooftops.
- The Restaurant: A Mediterranean cuisine space that maintains nods to traditional Swedish gastronomy.
View this program: with abstractssession overviewtalk overview
| 09:00 | Automatic Generation of Portable RVV Micro-Kernels through a Hybrid MLIR--xDSL Compilation Framework (abstract) PRESENTER: Jie Lei |
| 09:30 | Leveraging Dynamic Resource Management for Energy Efficiency: To Speed-up or To Green-up (abstract) |
| 10:00 | Exploiting Virtual Topology for Load- and Communication-Aware Neighborhood Collectives (abstract) |
| 10:30 | Phoning Home: Repatriating output from cloud jobs with a user-level, provider-agnostic data transfer tool (abstract) |
Tasks Based Hybrid Workflows for Emergent Technologies
The Barcelona Supercomputing Center hosts a twin installation based on transmon qubits, superconducting technology funded by Spain. Another annealing system co-funded by EuroHPC JU and Spain is also expected. Significant research and development is done by BSC researchers from the Quantic groups in topics related to quantum computing, which are closer to its physics nature. Complementary to this research, and in collaboration with our colleagues, the Workflows and Distributed Computing group is doing research activities related to software aspects related to the integration of HPC + Quantum Computing (QC).
The talk will describe the steps taken towards integration of HPC and QC at the BSC. We will describe how a traditional workflow environment is being extended to support hybrid HPC+quantum computing workflows.The current solution enables the hybrid execution of workflows in CPUS, GPUs and Quantum systems (onsite and through cloud access). The talk will present these topics and our plans for the near future.
| 12:00 | GPU-Efficient XAI Uncertainty Estimation for nnUNet (abstract) PRESENTER: Máximo Rodríguez Herrero |
| 12:30 | Design and Evaluation of Continuous Synthetic Data Pipelines for Vision Pre-training (abstract) PRESENTER: Ferran Soler |
| 13:00 | Parallelizing Deep Learning: A Unified Study of CPU and GPU Partitioning Strategies, Bottlenecks, and Design Principles (abstract) |
| 13:30 | Evaluating Runtime Prediction and Conservative Walltime Policies in HPC Systems (abstract) PRESENTER: Hector Romero Ugalde |