← Back to KHAO

Apple ·

Accelerating Text-to-Video Generation with Calibrated Sparse Attention

2 min read

Compiled by KHAO Editorial — aggregated from 1 source + 1 reference discovered via search. See llms.txt for citation guidance.

✓ KHAO Verified

Bottom banner.

Accelerating Text-to-Video Generation with Calibrated Sparse Attention.

Key facts

Summary

Authors Shai Yehezkel†**, Shahar Yadin, Noam Elata, Yaron Ostrovsky-Berman, Bahjat Kawar. Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled models at various resolutions show that CalibAtt achieves up to 1.58× end-to-end speedup, outperforming existing training-free methods while maintaining video generation quality and text-video alignment.

Read full article at Apple Machine Learning →

#Apple