After the Step 5.5 self-check passes, show a storyboard-mode choice card before producing any storyboard artifact:
Store the chosen storyboard mode in the Project Brief and reuse it in Step 7, Step 9, and the Regeneration discipline.
Generate one canvas text node named <title> text storyboards (one document for the whole short). This document is the authoritative rendering reference for Step 7 even when pencil images are also produced. The structure mirrors the half-narrated-drama storyboard — every shot is a section in the same document, so the user can read cross-shot continuity without node-hopping.
Document top matter (header block at the top of the document):
shot-table self-check: passed at <timestamp>).S01, S02, …) so the user can jump.Per-shot section structure (one ## heading per shot, in shot order). Every section is mandatory to contain these fields, in this order — direct adaptation of the half-narrated-drama storyboard:
S<N> / <duration>s (e.g. S03 / 6s).setup / visual-joke / reversal / reveal / callback / suspense / tender / chase / expression-beat / climax. Used by the per-episode hook distribution self-check.Fixed landmarks — named landmarks and their screen-relative positions (e.g. door-frame: right third, kitchen-island: center bottom).Character positions (camera view) — for every on-screen character, screen-relative position, facing direction, and initial pose.Exited character status — characters who were on screen in the previous shot but not in this one, with their off-screen position and reason.Lighting baseline — inherited key/fill/rim direction from the scene card, plus per-shot modifier.Continuity from S(N-1) — one or two sentences referencing the previous shot’s ending state.Continuity to S(N+1) — one sentence setting up the next shot’s opening.[char:角色名-01] [char:角色名-02] ... [scene:场景名] [hook: visual-joke] — exact character card names, scene card name, hook type. These are storyboard-only reference markers; the video model strips them at render time.Timecode — e.g. 0–1s.Pose + Expression — concrete body posture, silhouette, key prop grip, eye-line, facial expression path; for elastic beats, explicitly call out squash / stretch / anticipation / overshoot. This is the largest section per panel and is what the video model reads as the visual beat.Camera — shot size, camera movement (push / pull / pan / tilt / handheld-shake / locked / orbit), Dutch angle note when applicable.Audio + Anchor — audio cue (♪ narration: ... / dialogue: ... / SFX: ... / silent) and spatial anchor note (door-frame: right third / Mia: center midground facing camera).narrator-mouth-closed: true; for on-screen dialogue, mark mouth-open: speaker and describe expression path / eye-line / body-action changes.2.0–2.5s only when the beat is the hook of the shot.[0-1s] Mia (L, mid) door-frame (R)
──kneels, hands on apple basket──
cam: low push-in, locked
audio: silent | anchor: basket center-bottom
[1-2s] ...Per-panel four-quadrant content above, not the ASCII.[BEAT] after the panel timecode.[HANDOFF → ...] with a short label such as [HANDOFF → S04 opening].Per-shot section template (copy-paste skeleton, valid for any shot):
## S03 / 6s — Title: 奶奶把苹果筐递给 Mia
- **Hook type**: reveal
- **Scene & characters**: scene:kitchen | char:Mia, char:Grandma
- **Spatial anchor card**:
- Fixed landmarks: door-frame (right third), kitchen-island (center bottom)
- Character positions: Mia (L, midground, facing camera) | Grandma (R, foreground, facing Mia)
- Exited character status: —
- Lighting baseline: warm overhead key + cool bounce right
- **Continuity from S02**: 奶奶弯下腰从中岛拿起苹果筐
- **Continuity to S04**: Mia 接住筐转身,门铃响起
- **Double-binding**: [char:Mia] [char:Grandma] [scene:kitchen] [hook:reveal]
### Per-panel four-quadrant content
#### 0–1s
- Pose + Expression: 奶奶弯腰双手持筐;Mia 左侧站姿,眼神好奇
- Camera: locked medium shot, eye-level
- Audio + Anchor: silent | Mia: L midground | basket: center bottom
- Performance: [BEAT]
#### 1–2s
- Pose + Expression: 奶奶手臂伸向 Mia,筐倾斜;Mia 双手前伸准备接
- Camera: locked medium shot, eye-level
- Audio + Anchor: ♪ SFX: basket rustle | anchor: door-frame: right third
- Performance: [HANDOFF → S04 opening]
#### 2–3s
...
### ASCII layout (optional)
[0-1s] Grandma (R, fg) door-frame (R, bg)
──lifts basket── Mia (L, mid)
cam: locked | silent
[1-2s] ...
After all sections are written, place the document on canvas and move directly to Step 7. Do not call any image generation model in default mode.
The default single-document form is optimized for reading and cross-shot continuity. When the user flags a specific shot for heavy iteration (typically climax / chase / slapstick beats where the per-panel content needs many rounds of revision), extract that section into a standalone text node so iteration is localized:
<title> S05 text storyboard (extracted).## S05 section from the document into the new node.## S05 section with a one-line placeholder: > S05 — extracted to standalone node (see<title> S05 text storyboard (extracted)).The extraction mechanism exists because independent nodes are best used by need, not by default — but they remain available whenever iteration pressure is high on a specific shot.
If the user picked the visualization mode in the storyboard-mode choice card, ALSO produce one multi-panel pencil storyboard image per table row on top of the text storyboards document. The text storyboards document remains the authoritative rendering reference; the pencil images are human-review-only.
For each pencil image storyboard:
[char:角色名-01] [char:角色名-02] ... — exact character card names used in this row.[scene:场景名] — exact scene card name.[shot: S03] [dur: 6s] [hook: visual-joke] — shot ID, duration, and hook type.0–1s).♪ narration: "I knew it." / SFX: door creak / silent) and anchor note (e.g. door-frame: right third).0–1s, 1–2s, and show the corresponding pose, expression, action, camera movement, prop position, SFX cue, and continuity handoff.After all text storyboard sections (and pencil images, if visualization mode is on) are produced, place them on canvas in shot order, group them as:
<title> text storyboards (default mode, single document), OR<title> text storyboards + multi-panel pencil storyboards (visualization mode, group the text document and the pencil images separately because pencil images contain double-binding labels, ASCII labels, and shot numbers that the text document does not).Show a user choice card:
If a pencil image storyboard cannot be produced at the required quality (e.g. layout collapses, labels illegible, panels merged, character inconsistency), apply the following escalation before asking the user:
[char:…] [scene:…] [shot:…] labels, and the per-panel content rules.♪ mark) to reduce text load; this usually fixes illegible labels without losing the visual beat.In default text mode this whole fallback is unnecessary — text storyboards fail only when the model cannot produce coherent structured text, in which case return to Step 5 to revise the table row.