SCIC: Scope- and Codebook-Aware Instruction Conditioning
for Speaker-Adapted Expressive TTS

Longyu Lu, Zongwei Du, Mengtao Xing, Zhuoqun Liu, Zifan Guan, Meiguang Jin*, Junfeng Ma

TaoLive-AIGC Team, Taobao & Tmall Group of Alibaba

* Corresponding author

Abstract

Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline Prosody Control, where each Pitch, Energy, or Speed instruction targets a clause relative to the preceding clause from the same speaker, while Pause uses an absolute duration interval. In codec-based TTS, Speed and Pause affect sequence length, whereas Pitch and Energy rely on residual codebooks. By analyzing Qwen3-TTS RVQ codebooks, we find that Energy concentrates in early residual codebooks, whereas Pitch accumulates across a deeper prefix. We therefore propose Scope- and Codebook-Aware Instruction Conditioning (SCIC), combining a Temporal Instruction Router for frame-level tag activation with Tag-Specific Codebook Weighting over residual codebooks. SCIC improves speaker-relative Pitch and Energy control over standard instruction fine-tuning using text-token tags. We further apply multi-reward GDPO post-training to jointly optimize control and quality, improving control accuracy while preserving CER and speaker similarity. In long-form synthesis, SCIC produces a more distinct paragraph-level expressive hierarchy than speaker-adapted SFT without instructions.

Architecture

SCIC architecture

SCIC consists of four main components:

Single-Instruction Samples

Each sample contains one inline instruction tag embedded in the text. The tag controls the following clause relative to the preceding clause.

Multi-Instruction Samples

Long-form samples with multiple inline tags, combining Pitch, Energy, Speed, and Pause instructions within a single paragraph.