| ||||
| ||||
![]() Title:SMEAtten: Fast and Memory-Efficient Outer Product-Based Attention on ARMv9 CPUs with SME Conference:Euro-Par 2026 Tags:ARMv9, Attention and Scalable Matrix Extension Abstract: Transformer-based models are widely used in modern artificial intelligence, and the attention mechanism is a major determinant of their runtime efficiency. To accelerate matrix-intensive workloads, such as attention, ARMv9 introduces the Scalable Matrix Extension (SME) to enhance matrix processing capabilities on ARM CPUs. However, efficiently accelerating attention with SME remains challenging because of suboptimal compute-unit utilization, inefficient memory access, and limited task-level parallelism. To address these challenges, we present SMEAtten, a fast and memory-efficient attention design for ARMv9 CPUs with SME. SMEAtten incorporates three key techniques: throughput-driven interleaved matrix-vector attention kernels, an SME-adapted data layout and access scheme based on blocking, packing, buffering, and access-compute overlap, and inter-task parallelism exploitation. Experimental results show that SMEAtten delivers an average speedup of 13.62× over state-of-the-art baselines and achieves up to 56.09× acceleration in PyTorch integration evaluations. SMEAtten also supports efficient execution on different ARMv9 platforms, including LX2 and Apple M4. SMEAtten: Fast and Memory-Efficient Outer Product-Based Attention on ARMv9 CPUs with SME ![]() SMEAtten: Fast and Memory-Efficient Outer Product-Based Attention on ARMv9 CPUs with SME | ||||
| Copyright © 2002 – 2026 EasyChair |
