Apple · Apple Machine Learning
Accelerating Text-to-Video Generation with Calibrated Sparse Attention
Compiled by KHAO Editorial — aggregated from 1 source + 1 reference discovered via search. See llms.txt for citation guidance.
✓ KHAO Verified
Accelerating Text-to-Video Generation with Calibrated Sparse Attention.
Key facts
- Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled models at various resolutions show that CalibAtt achieves up to 1.58× end-to-end speedup, outperforming existing training-free
- Authors Shai Yehezkel†**, Shahar Yadin, Noam Elata, Yaron Ostrovsky-Berman, Bahjat Kawar
- Accelerating Text-to-Video Generation with Calibrated Sparse Attention
- Recent diffusion models enable high-quality video generation, but suffer from slow runtimes
Summary
Authors Shai Yehezkel†**, Shahar Yadin, Noam Elata, Yaron Ostrovsky-Berman, Bahjat Kawar. Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled models at various resolutions show that CalibAtt achieves up to 1.58× end-to-end speedup, outperforming existing training-free methods while maintaining video generation quality and text-video alignment.