AptAvatar:Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Hengyuan ZhangJingna SunMeiguang JinJunfeng Ma

Taobao & Tmall Group of Alibaba

† Equal contribution✉ Corresponding author

A powerful fine-tuned multi-step generator distilled to two inference steps—while retaining high fidelity, expressive motion, and long-horizon identity consistency.

14Bfine-tuned generator
720phigh-resolution video
2NFEs at inference
60×inference speedup
166-SECOND OUTPUTExtended long-video demo
2-NFE generation
Identity-consistent
VIDEO DEMOS

Short and long videos

Cartoon short videos, real-person short videos, and minute-level long videos.

SHORT VIDEO · CARTOON

Cartoon generation

10 short-video demos

01
Wood Puppet15 seconds · stylized avatar
02
Cel-Shaded Anime15 seconds · stylized avatar
03
Dollhouse15 seconds · stylized avatar
04
Clay Stop-Motion15 seconds · stylized avatar
05
3D Ink Wash15 seconds · stylized avatar
06
Pastel Diorama15 seconds · stylized avatar
07
Ink-Wash Toon15 seconds · stylized avatar
08
Porcelain Figurine15 seconds · stylized avatar
09
Art Nouveau15 seconds · stylized avatar
10
Cyberpunk Anime15 seconds · stylized avatar
SHORT VIDEO · REAL PERSON

Expressive real-person generation

25 short-video demos

01
Real-person demo 0115 seconds · 704×1280
02
Real-person demo 0215 seconds · 704×1280
03
Real-person demo 0315 seconds · 704×1280
04
Real-person demo 0415 seconds · 704×1280
05
Real-person demo 0515 seconds · 704×1280
06
Real-person demo 0615 seconds · 704×1280
07
Real-person demo 0715 seconds · 704×1280
08
Real-person demo 0815 seconds · 704×1280
09
Real-person demo 0915 seconds · 704×1280
10
Real-person demo 1015 seconds · 704×1280
11
Real-person demo 1115 seconds · 704×1280
12
Real-person demo 1215 seconds · 704×1280
13
Real-person demo 1315 seconds · 704×1280
14
Real-person demo 1415 seconds · 704×1280
15
Real-person demo 1515 seconds · 704×1280
16
Real-person demo 1615 seconds · 704×1280
17
Real-person demo 1715 seconds · 704×1280
18
Real-person demo 1815 seconds · 704×1280
19
Real-person demo 1915 seconds · 704×1280
20
Real-person demo 2015 seconds · 704×1280
21
Real-person demo 2115 seconds · 704×1280
22
Real-person demo 2215 seconds · 704×1280
23
Real-person demo 2315 seconds · 704×1280
24
Real-person demo 2415 seconds · 704×1280
25
Real-person demo 2515 seconds · 704×1280
LONG VIDEO

Long-video generation

4 minute-level demos

01
Long-form demo 0148.5 seconds · 704×1280
02
Long-form demo 0248.4 seconds · 704×1280
03
Long-form demo 0340.6 seconds · 704×1280
04
Long-form demo 0448.4 seconds · 704×1280
MOTIVATION

Production-ready generation requires both speed and long-video quality.

Fast inference must be balanced with motion quality and identity consistency over long videos.

AptAvatar teaser showing reference, audio, and text inputs; vivid short-video motion; and identity-consistent long-video generation
AptAvatar accepts a reference image, speech audio, and fine-grained text prompts, then produces expressive short-form motion and stable long-form video with only two inference steps.
01

Extreme acceleration, no capacity cut

Existing speedups often causalize attention, shrink the model, or shorten its temporal window. AptAvatar keeps the 14B bidirectional backbone and compresses its denoising trajectory instead.

02

Long-form generation without drift

Chunk-wise generation must condition on its own imperfect history. Small errors accumulate into identity drift, degraded hands, and reduced motion expressiveness unless training reproduces that deployment reality.

METHOD

Anchor the endpoint. Replay the history.

Two complementary designs make aggressive two-step distillation and stable autoregressive rollout work together.

AptAvatar framework with Endpoint-Anchored Distribution Distillation and Self-Generated History Replay
01

Endpoint-Anchored Distribution Distillation

A frozen four-step bridge generator defines an attainable trajectory endpoint. A dedicated Anchor Score Estimator turns that fixed endpoint into stable guidance for the evolving two-step student.

EADD
02

Self-Generated History Replay

Cached outputs from earlier generator checkpoints become history conditions during training, approximating inference-time self-conditioning without costly recurrent online rollouts.

SGHR
QUANTITATIVE

Two steps, competitive across both benchmarks.

AptAvatar uses the fewest NFEs and leads most fidelity, consistency, and synchronization metrics reported in the paper.

15s
Short-video benchmark100 videos
BG-C ↑
96.844
Temporal-F ↑
99.456
Sync-C ↑
7.865
NFE ↓
2
60s
Minute-level benchmark20 videos
Subject-C ↑
98.895
BG-C ↑
96.489
Temporal-F ↑
99.623
Sync-C ↑
8.398
VS. SOTA

Comparison with state of the art

Short- and long-video comparisons are separated so motion quality and accumulated temporal degradation can be inspected independently.

SHORT VIDEO

15-second comparison

Same reference and audio across all five methods. Playback and seeking are synchronized.

OURS
AptAvatarOurs · 2 NFE
LiveAvatar4 NFE
SoulX-FlashTalk4 NFE
Wan2.2-S2V80 NFE
LongCat-Avatar 1.58 NFE
LONG VIDEO

Minute-level temporal comparison

Same reference and audio across all five methods. Playback and seeking are synchronized.

OURS
AptAvatarOurs · 2 NFE
LiveAvatar4 NFE
SoulX-FlashTalk4 NFE
Wan2.2-S2V80 NFE
LongCat-Avatar 1.58 NFE