The Missing Temporal Link — Temporal Context Routing
Script-driven audio-video generation: describe a clip as a JSON script whose shots and
events carry explicit time_range timings, and the model cuts, paces and syncs the audio to
that timeline. TCR routes each text token to the video/audio tokens that fall inside its time
interval via an additive bias on text cross-attention.
LTX-2.3 22B (fp8) + the TCR rank-128 LoRA + Gemma-3-12B-IT text encoder. paper · model · code
Script duration: 10.00s → 241 frames @ 24 fps
| # | type | time_range | content |
|---|---|---|---|
| 1 | SHOT_1 | 0.00–2.20s |
[PERSON_1] and [PERSON_2] are standing outside [SCENE_1]. [PERSON_1] turns away slightly. |
| 2 | SHOT_2 | 2.20–3.80s |
A close-up of [PERSON_1] looking at [PERSON_2] and speaking. |
| 3 | SHOT_3 | 3.80–5.20s |
A close-up of [PERSON_2] looking at [PERSON_1] and speaking. |
| 4 | SHOT_4 | 5.20–6.80s |
A close-up of [PERSON_1] smiling and nodding as [PERSON_2] continues to speak off-screen. |
| 5 | SHOT_5 | 6.80–8.50s |
Viewed from inside the house, [PERSON_3] appears in the doorway, looking surprised. |
| 6 | SHOT_6 | 8.50–10.00s |
A close-up of [PERSON_2] turning his head with a surprised expression. [PERSON_1] is partially visible in t… |
| 1 | DIALOGUE_1 | 1.50–2.80s |
PERSON_1: 没有 |
| 2 | DIALOGUE_2 | 2.80–4.40s |
PERSON_1: 我们先回去吧 |
| 3 | DIALOGUE_3 | 4.40–5.60s |
PERSON_2: 那先上车吧 |
| 4 | DIALOGUE_4 | 5.60–5.90s |
PERSON_2: 你别着凉 |
Global audio: Quiet outdoor ambience with clear dialogue.
The number of frames is derived from the script itself (the latest time_range end time), so the generated clip always matches the timeline you wrote. Max 10s @ 24 fps.
Examples