Join our Discord to engage in discussions with the community! If you have any questions, run into issues, or are interested in contributing, don't hesitate to reach out!
News
π₯ [2026/09] β‘ SANA-Video 2.0 5B 4-Step Preview is released! The DMD preview generates 720p videos in four denoising steps and supports 5-second and 8-second outputs. See Online Demo | 4-Step Weights | Doc | Project.
π₯ [2026/08] π¬ SANA-Video 2.0 training, inference, model architecture, and 5B 720p checkpoint are released! The 8-second model supports both text-to-video and text-image-to-video generation, with hybrid linear/softmax attention and Attention Residuals. See Online Demo | Project | Doc | Model Zoo | Weights.
π₯ [2026/08] β‘ Sol Engine: Day-One MiniMax-H3 Acceleration is available! The 33B omni-modal audio+video DiT runs 3.95Γ faster on GB200, reached in 4.5 hours of optimization, and up to 4.52Γ on hardware that sits on a desk β 3.92Γ on DGX Spark, 4.52Γ on GeForce RTX 5090 β with no distillation, no LoRA, and no calibration pass. See GB200 Blog | On-Device Blog.
π₯ [2026/07] π SANA-Streaming training is released! Includes bidirectional and distillation training. See Doc.
π₯ [2026/07] π SANA-WM Stage-1 training is released! Includes bidirectional, chunk-causal, and distillation training. See Doc.
π₯ [2026/06] π¬ SANA-Streaming: 2B Model for Real-time Streaming Editing is released! Supports 720p, 1-min video editing. A pioneer work for streaming editing. See Project | Doc | Paper | Reactor Demo.
π₯ [2026/05] π SANA-WM: 2.6B Controllable World Model is released! Supports 720p, 1-min video generation with 6-DoF camera control. A new baseline for World Modeling and Embodied AI. See Project | Doc | Paper | Reactor Demo.
π₯ [2026/04] β‘ Sol-RL: NVFP4 Rollout, BF16 Training RL is available! All training recipes for SANA, FLUX.1, and SD3.5-L, together with bundled post-training datasets, are released. See Sol-RL doc | Page | Paper.
π₯ [2026/03] πΊ SANA-Video 720p model with LTX-VAE is released. Use it with LTX2 Refiner to upscale the videos to 2K resolution! See Model Zoo, SANA-Video doc and Blog about refiner.
π₯ [2026/03] πͺ Post Training Infra: SANA Γ Cosmos-RL β We partner with Cosmos-RL to provide a complete RL infrastructure for SANA. You can now post-train (SFT/RL) SANA-Image and SANA-Video with state-of-the-art algorithms (e.g. Diffusion-NFT, Flow-GRPO), preset configs, reward services, and flexible datasets. See SANA on Cosmos-RL and our Cosmos-RL integration doc.
π₯ [2026/02] π SANA is now supported in SGLang! High-performance serving with OpenAI-compatible API. [Guidance]
π₯ [2026/01/26] SANA-Video is accepted as Oral by ICLR-2026. πππ
π₯ [2025/12/09] π¬ LongSANA: 27FPS real-time minute-length video generation model, training and inference code are all released. Thanks to LongLive Team. Refer to: [Train] | [Test] | [Weight]
π₯ [2025/11/24] πͺΆ Blog: how Causal Linear Attention unlocks infinite context for LLMs and long video generation.
π₯ [2025/11/9] π¬ Introduction video shows how Block Causal Linear Attention and Causal Mix-FFN work?
π₯ [2025/10/27] πΊSANA-Video is released. [README] | [Weights] support Text-to-Video, TextImage-to-Video.
π₯ [2025/10/13] πΊSANA-Video is coming, 1). a 5s Linear DiT Video model, and 2). real-time minute-length video generation (with LongLive). [paper] | [Page]
β [2025/3/14] πSANA-Sprint is coming out! π A new one/few-step generator of Sana. 0.1s per 1024px image on H100, 0.3s on RTX 4090. Find out more details: [Page] | [Arxiv]. Code is coming very soon along with diffusers
β [2025/1/24] DCAE-1.1 is released, better reconstruction quality. [Model][diffusers]
β [2025/1/23] Sana is accepted as Oral by ICLR-2025. πππ
β [2025/1/12] DC-AE tiling makes Sana-4K inferences 4096x4096px images within 22GB GPU memory. With model offload and 8bit/4bit quantize. The 4K Sana run within 8GB GPU VRAM. [Guidance]
β [2025/1/11] Sana code-base license changed to Apache 2.0.
β [2025/1/10] Inference Sana with 8bit quantization.[Guidance]
β [2025/1/8] 1.6B 4K resolution Sana models are released: [BF16 pth] or [BF16 diffusers]. π Get your 4096x4096 resolution images within 20 seconds! Find more samples in Sana page. Thanks SUPIR for their wonderful work and support.
β [2025/1/2] Bug in the diffusers pipeline is solved. Solved PR
β [2024/12] 1.6B 2K resolution Sana models are released: [BF16 pth] or [BF16 diffusers]. π Get your 2K resolution images within 4 seconds! Find more samples in Sana page. Thanks SUPIR for their wonderful work and support.
β [2024/12] diffusers supports Sana-LoRA fine-tuning! Sana-LoRA's training and convergence speed is super fast. [Guidance] or [diffusers docs].
β [2024/12] diffusers has Sana! All Sana models in diffusers safetensors are released and diffusers pipeline SanaPipeline, SanaPAGPipeline, DPMSolverMultistepScheduler(with FlowMatching) are all supported now. We prepare a Model Card for you to choose.
β [2024/12] 1.6B BF16 Sana model is released for stable fine-tuning.
SANA-Video 2.0: 5B and 14B text-to-video/text-image-to-video architectures with hybrid linear/softmax attention and Attention Residuals. Try the 5B 720p 4-step preview, or download the 50-step and 4-step preview checkpoints; the 14B config and checkpoint are not included yet.
Sol-RL: NVFP4 Rollout, BF16 Training RL achieves 4.64Γ faster convergence.
SANA-WM: 2.6B parameter controllable world model, generating 720p, 1-minute video worlds with 6-DoF camera control.
SANA-Streaming: 2B real-time streaming video-to-video editing for 720p, minute-scale videos.
Key Techniques:
Linear Attention: Replace vanilla attention in DiT with linear attention for efficiency at high resolutions.
DC-AE: 32Γ image compression (vs. traditional 8Γ) to reduce latent tokens.
Decoder-only Text Encoder: Modern decoder-only LLM with in-context learning for better text-image alignment.
Block Causal Linear Attention & Causal Mix-FFN: Efficient attention and feedforward for long video generation.
Hybrid Attention & Attention Residuals: Combine gated linear attention with periodic softmax anchors and shared depth-wise residual aggregation.
Flow-DPM-Solver: Reduce sampling steps with efficient training and sampling.
sCM Distillation: One/few-step generation with continuous-time consistency distillation.
Sol-RL: Low precision(NVFP4) rollout selection, high precesion(BF16) optimization for faster RL training.
Controllable World Modeling: Efficient long-context modeling and camera trajectory control for consistent world generation.
Streaming Video Editing: Real-time long-form video-to-video editing with stable temporal consistency.
In summary, SANA is a fully open-source framework integrating efficient training, fast inference, and flexible deployment for both image and video generation. Deployable on laptop GPUs with < 8GB VRAM via 4-bit quantization.
Quick Start
git clone https://github.com/NVlabs/Sana.git
cd Sana && ./environment_setup.sh sana
SANA-Video 2.0 5B release demo
This sample was generated from the public 5B checkpoint with seed 4. The result
contains 193 frames at 24 FPS in a 1280 Γ 736 bucket (8.04 seconds).
Prompt: In a cozy, vintage room adorned with floral wallpaper, a cartoon
rooster sits comfortably in a floral-patterned armchair, sipping from a bottle
of beer. The rooster, with its vibrant red comb and wattle, displays a range of
expressionsβsmiling, nodding, and opening its beak wide in a cheerful manner.
The setting includes wooden furniture and another beer bottle on the table,
adding to the relaxed atmosphere. The camera captures the rooster from a
close-up angle, emphasizing its animated movements and lively demeanor.
Run the exact release command used for the video above:
The online preview supports 5-second (81 frames at 16 FPS) and 8-second
(193 frames at 24 FPS) outputs, and defaults to RL LoRA scale 0.7. Run the
T2V-only full DMD checkpoint at its native RL scale 1.0 with the 5-second
profile using:
The verified online seed-4 preview below uses the same prompt, temporal profile,
and motion score, with the Space default RL LoRA scale 0.7 (1280 Γ 736,
81 frames, 16 FPS, 5.06 seconds). The model card records its exact API call.
Thanks Paper2Video for generating Jeason presenting SANAπ. Refer to Paper2Video for more details.
Contribution
Thanks go to these wonderful contributors:
π Star History
π BibTeX
@misc{xie2024sana,
title={Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer},
author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
year={2024},
eprint={2410.10629},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2410.10629},
}
Click to expand all BibTeX citations
@misc{xie2025sana,
title={SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer},
author={Xie, Enze and Chen, Junsong and Zhao, Yuyang rectangle and Yu, Jincheng and Zhu, Ligeng and Lin, Yujun and Zhang, Zhekai and Li, Muyang and Chen, Junyu and Cai, Han and others},
year={2025},
eprint={2501.18427},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2501.18427},
}
@misc{chen2025sanasprint,
title={SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation},
author={Junsong Chen and Shuchen Xue and Yuyang Zhao and Jincheng Yu graves and Sayak Paul and Junyu Chen and Han Cai and Song Han and Enze Xie},
year={2025},
eprint={2503.09641},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.09641},
}
@misc{chen2025sanavideo,
title={SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer},
author={Chen, Junsong and Zhao, Yuyang and Yu, Jincheng and Chu, Ruihang and Chen, Junyu and Yang, Shuai and Wang, Xianbang and Pan, Yicheng and Zhou, Daquan and Ling, Huan and others},
year={2025},
eprint={2509.24695},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.24695},
}
@misc{li2026fp4,
title={FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling},
author={Li, Yitong and Chen, Junsong and Xue, Shuchen and Zeren, Pengcuo and Fu, Siyuan and Yang, Dinghao and Tang, Yangyang and Bai, Junjie and Luo, Ping and Han, Song and others},
year={2026}
eprint={2604.06916},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.06916},
}
@misc{zhu2026sanawm,
title={SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer},
author={Haoyi Zhu and Haozhe Liu and Yuyang Zhao and Tian Ye and Junsong Chen and Jincheng Yu and Tong He and Song Han and Enze Xie},
year={2026},
eprint={2605.15178},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.15178},
}
@misc{zhao2026sanastreamingrealtimestreamingvideo,
title={SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer},
author={Yuyang Zhao and Yicheng Pan and Qiyuan He and Jincheng Yu and Junsong Chen and Tian Ye and Haozhe Liu and Enze Xie and Song Han},
year={2026},
eprint={2605.30409},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.30409},
}