| ||||
| ||||
![]() Title:Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication Conference:Euro-Par 2026 Tags:DiT, Efficient Communication and Sequence Parallelism Abstract: Sequence Parallelism (SP), such as DeepSpeed-Ulysses, has emerged as the de facto standard for scaling high-resolution video generation for its superior memory efficiency in handling long contexts and reduced communication. However, the bandwidth-intensive All-to-All communication inherent in SP still becomes a critical bottleneck, severely limiting scaling efficiency. We propose Sparsh, a distributed video diffusion inference framework with predictive sparse communication. Unlike standard Ulysses-style sequence parallelism which indiscriminately transmits attention tensors in all diffusion steps, Sparsh eliminates the transmission of Key and Value tensors entirely in most steps by exploiting their temporal redundancy. Particularly, instead of remotely fetching KV tensors via All-to-All communication, Sparsh introduces a fidelity-aware predictive synthesis mechanism that reconstructs remote KV states locally. Further, Sparsh employs an online adaptive controller that dynamically adjusts prediction strategies based on real-time error feedback, ensuring that the communication reduction does not compromise generation quality. Experiments on multiple models demonstrate that Sparsh reduces communication latency by 50% in predictive steps and by 28% overall compared to state-of-the-art Ulysses baselines, offering a scalable solution for distributed high-resolution video generation. Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication ![]() Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication | ||||
| Copyright © 2002 – 2026 EasyChair |
