arxiv:2609.20744
Published on Sep 17
· Submitted by
taesiri on Sep 18
Upvote
47
Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.
View arXiv page View PDF Project page Add to collection
Community
4 days ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
-
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation (2026)
-
HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation (2026)
-
LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding (2026)
-
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention (2026)
-
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers (2026)
-
Self Gradient Forcing: Native Long Video Extrapolation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on HF Mirror checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
4 days ago
This comment has been hidden (marked as Low Quality)
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
· Sign up or log in to comment
Upvote
47
Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.20744 in a dataset README.md to link it from this page.
Spaces citing this paper 18
Browse 18 spaces citing this paper
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.