AptAvatar:Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars
† Equal contribution✉ Corresponding author
A powerful fine-tuned multi-step generator distilled to two inference steps—while retaining high fidelity, expressive motion, and long-horizon identity consistency.
Short and long videos
Cartoon short videos, real-person short videos, and minute-level long videos.
Cartoon generation
10 short-video demos
Expressive real-person generation
25 short-video demos
Long-video generation
4 minute-level demos
Production-ready generation requires both speed and long-video quality.
Fast inference must be balanced with motion quality and identity consistency over long videos.

Extreme acceleration, no capacity cut
Existing speedups often causalize attention, shrink the model, or shorten its temporal window. AptAvatar keeps the 14B bidirectional backbone and compresses its denoising trajectory instead.
Long-form generation without drift
Chunk-wise generation must condition on its own imperfect history. Small errors accumulate into identity drift, degraded hands, and reduced motion expressiveness unless training reproduces that deployment reality.
Anchor the endpoint. Replay the history.
Two complementary designs make aggressive two-step distillation and stable autoregressive rollout work together.
Endpoint-Anchored Distribution Distillation
A frozen four-step bridge generator defines an attainable trajectory endpoint. A dedicated Anchor Score Estimator turns that fixed endpoint into stable guidance for the evolving two-step student.
Self-Generated History Replay
Cached outputs from earlier generator checkpoints become history conditions during training, approximating inference-time self-conditioning without costly recurrent online rollouts.
Two steps, competitive across both benchmarks.
AptAvatar uses the fewest NFEs and leads most fidelity, consistency, and synchronization metrics reported in the paper.
- BG-C ↑
- 96.844
- Temporal-F ↑
- 99.456
- Sync-C ↑
- 7.865
- NFE ↓
- 2
- Subject-C ↑
- 98.895
- BG-C ↑
- 96.489
- Temporal-F ↑
- 99.623
- Sync-C ↑
- 8.398
Comparison with state of the art
Short- and long-video comparisons are separated so motion quality and accumulated temporal degradation can be inspected independently.
15-second comparison
Same reference and audio across all five methods. Playback and seeking are synchronized.
Minute-level temporal comparison
Same reference and audio across all five methods. Playback and seeking are synchronized.