# Faster dramatic speech and continuously new pictures

The main reference feels quicker through **dense spoken clauses, immediate character
responses, and a new useful view**. The previous film replaced that with a reflective
single narrator, long restarted compositions and repeated inward camera movement.
The next iteration needs a different dramatic and editorial language, not just playback
speed. These are source findings and intended directions; they do not approve the new film.

## Bound evidence

Main reference: `resilia-aged-garlic-family-cholesterol-story.mp4`, 325.681678 s,
SHA-256 `708cb005be27b73d43fa248bd8f003489098a48a43fcc59d6ce31bde841f7ff6`.
Previous film: `delivery/new-film/film-480p.mp4`, SHA-256
`12d63a16ead46b5cff0e23098767cf73b00c5ee8d4472198905fcc19f1a289c1`.

[Machine-readable findings](style-pace-measurements.json) bind the actual media, nine
bounded native-video analyses, actual source style frames, and frame observations.
[Picture measurements](picture-measurements.json) are reproduced by
`python3 scripts/measure_dynamic_picture.py`. Native analyses are retained without
rewriting their mistakes; source clocks and semantic counts remain model-assisted.

## Speech: the speed difference is almost twofold during explanation

| Actual excerpt | Orthographic words | Elapsed seconds | Words/minute |
|---|---:|---:|---:|
| Main reference, teaching, 77–105 s | 92 | 28 | 197.1 |
| Main reference, objection and answer, 160–190 s | 95 | 30 | 190.0 |
| Main reference, close, 284–325.56 s | 120 | 41.56 | 173.2 |
| Previous film, explanation, 109.36–157.68 s | 82 | 48.32 | 101.8 |
| Previous film, proof excerpt, 205.88–255.64 s | 78 | 49.76 | 94.1 |
| Previous entire accepted voice | 605 | 367.44 | 98.8 |

Source counts come from actual-media native transcripts. Previous excerpt counts also
match the accepted word-alignment ledger. A written numeral counts as one token;
hyphenated words and internal apostrophes remain together. Source transcript timing is
approximate. English/Ukrainian morphology differs, so these are useful pacing contrasts,
not acoustically exact cross-language articulation rates.

The opening reference averages only 102 WPM over 30 s; its search passage only 36 WPM.
Those are **not** slow explanatory monologues. The opening has clipped diagnosis/question
turns and silent processing; the search contains visible queries, realization, departure
and arrival. Averages hide how briskly speech and story move while they are active.
Do not fill every pause with narration. Make a pause hold a specific action or unresolved
reaction. The previous four-week section lasts 65.40 s and often narrates verification
rules such as identical lighting and unchanged face; those details belong in the report.

Native listening describes useful distinct performances:

- **Alarm, 0–13 s:** the doctor's declarative gravity meets the father's short rising
  disbelief and urgent question. Each line changes the response demanded of the listener.
- **Reassurance, about 18–25 s:** the daughter's vulnerable question meets the father's
  strained protective reply. The listener and the reassuring face are separate pictures.
- **Teaching, 77–105 s:** a grounded, calm voice still delivers dense clauses quickly;
  the child's short question interrupts explanation. Calm does not mean slow.
- **Objection, 160–190 s:** polite skepticism ends in a clear question; the helper picks
  it up with a quick corrective response, then explains. Do not insert ceremonial pauses
  before every sentence or repeat narrative tags describing who spoke.
- **Close, 284–325 s:** gratitude hands over to a direct practical address. The vocal role
  and energy change. Do not import its medical, scarcity or guarantee claims.

For the new Ukrainian voice, aim initially at **165–185 WPM**, with short 5–12-word
exchanges, intelligible consonants, contextual stress and audible emotional turns. A long
sentence may exceed that turn length, but it should carry new information or a question.
The number is an intended target; occurrence-level listening of the real faster voice
must decide whether it works. Do not use indiscriminate audio speed-up to hide flat acting.
Seedance lip-sync must reproduce those accepted words and emotions on the correct speaker.

## Pictures: coverage creates variation; constant push-ins erase it

The rerun detector finds **89 candidate cuts**, median interval **2.517 s**, mean **3.619 s**.
This is not a ground-truth shot count: the detector misses fades and visually similar edits.
There are four detected boundaries before 10 s and fourteen before 30 s, with an additional
visibly observed office-to-car transition around 15–16 s. The previous authored plan has
46 shots, median **7.78 s**, mean **7.99 s**, one boundary before 10 s and three before 30 s.
Native shot lists contain approximate timing and cannot replace the actual edit clock.

The [actual opening contact sheet](source-first30-contact.jpg) shows the useful grammar:
room master → father reaction reverse → doctor over shoulder → supported father → doctor
close → hand arriving on hand → daughter watching father → travelling exterior → daughter's
car view → father's reverse → family dashboard master → daughter close → family response
→ steering-wheel grip → rear-car view → new night bedroom. A hand insert or listener view
does emotional work while the ongoing line continues. The changing picture need not mean
a new room every two seconds.

The [source camera triplets](source-camera-triplets.jpg) establish bounded movement evidence:
the doctor at 4.1–7.8 s, hand insert at 11.8–13.2 s, helper at 80–84.5 s, and group master at
171.7–176 s keep their background geometry stable while bodies act. At 46–48 s the child
walks through a locked wide street composition. Subject walking is not camera tracking.
There is a different view by 50.5 s, so that triplet must not be described as one continuous
tracking shot. No whole-film optical-flow census was performed.

The [previous camera triplets](previous-camera-triplets.jpg) directly contradict the native
audits' generic “static” labels. F01a narrows a two-shot to the daughter; F01b restarts the
same setup and narrows toward the mother; F06a narrows a seated search to phone/head; F11a
narrows the patch test to the hand/forearm. Background scale/crop changes support actual
inward framing, rather than only the actor leaning. Do not count every clip as a push-in:
F06b and F20b show other behavior. The failure is a demonstrated repeated recipe.

## A different starting frame means a different composition

The previous 46 jobs use **19 unique storyboard hashes** with **26 adjacent same-board
pairs**. Actual entry sheets [1](previous-entries-1.jpg), [2](previous-entries-2.jpg),
[3](previous-entries-3.jpg) confirm repeated layouts with small pose changes. This is not
merely a bookkeeping complaint. Several beginning pairs show the same table, same axis,
same scale and same actor placements before another inward move. The eight proof entries
repeat frontal portrait grammar for more than a minute.

The new shot contract should separate **identity reference** from **scene composition**.
Each commissioned shot gets its own board describing axis, scale, foreground, actor
placement, ongoing action and immediate question. A new image hash or expression alone
does not prove new composition. Compare actual generated entry frames side by side.
Never split a long beat into two new jobs that restart the same board.

Keep proof measurable without creating eight restarted portraits: use one concise,
continuous four-state comparison. Four distinct state references, clear week labels,
matched concern region and short state dissolves preserve the passport. First use stays
at baseline. The comparison is an intentional controlled sequence, not a pretext to reuse
every dramatic opening frame. Visual verification should distinguish normal mature anatomy
from the declared dry/crepey concern; the native prior-proof output incorrectly called
week three a final result.

## Replace the visual style with the source's dimensional drama

[Four actual full-resolution source frames](style-frames/manifest.json) were extracted
and visually inspected for conditioning. They cover speaker reverse, listener reverse,
group master and bright outdoor action. Their role is style evidence only: exclude source
cast, child age, location, captions, packet and claims from the new content.

- Give faces a directional key, readable eyes and real shadow depth; retain texture and
  plausible highlights. The source has warm practical kitchens and cooler exterior/office
  spaces, not a universal dark filter.
- Compose real conversational axes: foreground shoulder or hand, off-center speaker,
  consistent lateral eyeline, listener reverse, then a wider spatial view when needed.
  Lens/depth impressions are observational; no actual focal length is known.
- Let the prop affect behavior: turn a phone toward someone, examine a label, grip a mug,
  put a hand on the correct person's hand. A product floating in a decorative tableau
  does not advance the scene.
- Use present-day Ukrainian rooms/courtyards through specific materials, architecture,
  clothes and ordinary action. Appeal can come from warm wood, clear window light, layered
  greenery and an engaging silhouette; neither forced gloom nor saturation is required.
- Use mostly clean cuts with distinct scale/angle. Reserve a motivated lateral follow,
  pan/reveal, high/low action insert or occasional deliberate inward move for its purpose.
  Do not replace one repetitive move with an arbitrary rotation of flashy effects.

## Callable transfer and next-film checks

**Function:** `compress_story(viewer_concern, affected_person, desired_activity,
credible_option, objection, valid_use, visible_change, local_world, action_destination)`.
Expose a concrete conflict through speech/action → answer it with a character choice →
make the option answer a real objection → show valid first use → compress elapsed change
into inspectable states → restore the desired activity → give one clear next action.
Each scene exits when knowledge, agency or visible state has changed. Family roles,
cosmetic weeks and clinical stakes are adapter choices, not universal invariants.

The previous opening has two annotated major turns over 32.44 s versus four in the source's
30 s; source search has five in 30 s. These are semantic annotations, not objective
retention scores. The teaching reference has only two major turns in 28 s despite fast
speech, which is why copying its entire exposition would not alone accelerate the story.

For the new iteration, verify actual media against these intended checks:

1. Conflict, specific concern and a decision appear within roughly the first 10–15 s.
   The complete story remains intact; shorter is useful only if the causal turns survive.
2. Typical useful views last about 1.5–3.5 s, with justified 4–6 s spoken holds. Cuts react
   to words, gestures, reveals or listener changes. Do not cut off dialogue to meet a quota.
3. New entry compositions and varied scale/axis are visible in a final entry contact sheet.
   There are no adjacent jobs restarting one scene board and no blanket push-in instruction.
4. Correct speaker, stable voice, Ukrainian pronunciation, emotional contour and actual
   mouth synchronization pass listening/moving-video review. Offscreen continuation over a
   listener or insert is deliberate and preserves the full sentence.
5. Concern ownership, first-use baseline, distinct weeks, resolved target and clothing
   changes still pass actual-pixel review. Pace does not excuse passport failures.
6. The same new iteration's report publishes the film, pacing comparison, source style
   evidence, entry frames, lip-sync checks and rejected attempts, with hosted playback.

No audience test, clinical verification, human linguistic certification, or measured
attention lift is claimed. Automated and assistant evidence has explicit scope; camera
labels and shot timing that contradict actual frames were corrected above.
