The Missing Temporal Link — Temporal Context Routing

Script-driven audio-video generation: describe a clip as a JSON script whose shots and events carry explicit time_range timings, and the model cuts, paces and syncs the audio to that timeline. TCR routes each text token to the video/audio tokens that fall inside its time interval via an additive bias on text cross-attention.

LTX-2.3 22B (fp8) + the TCR rank-128 LoRA + Gemma-3-12B-IT text encoder. paper · model · code

Script duration: 10.00s → 241 frames @ 24 fps

# type time_range content
1 SHOT_1 0.00–2.20s [PERSON_1] and [PERSON_2] are standing outside [SCENE_1]. [PERSON_1] turns away slightly.
2 SHOT_2 2.20–3.80s A close-up of [PERSON_1] looking at [PERSON_2] and speaking.
3 SHOT_3 3.80–5.20s A close-up of [PERSON_2] looking at [PERSON_1] and speaking.
4 SHOT_4 5.20–6.80s A close-up of [PERSON_1] smiling and nodding as [PERSON_2] continues to speak off-screen.
5 SHOT_5 6.80–8.50s Viewed from inside the house, [PERSON_3] appears in the doorway, looking surprised.
6 SHOT_6 8.50–10.00s A close-up of [PERSON_2] turning his head with a surprised expression. [PERSON_1] is partially visible in t…
1 DIALOGUE_1 1.50–2.80s PERSON_1: 没有
2 DIALOGUE_2 2.80–4.40s PERSON_1: 我们先回去吧
3 DIALOGUE_3 4.40–5.60s PERSON_2: 那先上车吧
4 DIALOGUE_4 5.60–5.90s PERSON_2: 你别着凉

Global audio: Quiet outdoor ambience with clear dialogue.

The number of frames is derived from the script itself (the latest time_range end time), so the generated clip always matches the timeline you wrote. Max 10s @ 24 fps.

Examples