ByteDance shipped 2.5 after 2.0. The version number tells you it is newer. It does not tell you it is better, and on two axes it is measurably worse.
Renders 34Tracks 13Spend 2087 / 1500 crRe-rolls 0Generated with Higgsfield
Made with
Every frame on this page was generated with Higgsfield. Seedance 2.5 and 2.0 both run on Higgsfield, which is where all twenty renders on this page were made — same account, same session, costs preflighted per call.
Most model comparisons are two nice-looking clips side by side. That is a taste test, not a test.
One world, one cast, one look. A rain-wet container dock past midnight — four sodium floods
high and behind, one cold blue work lamp low and camera-left, palette locked at 60 wet
asphalt / 25 sodium orange / 15 cold blue.
That CARD was written once and reused verbatim across every live-action track.
Then seven TAKE blocks — action, K-pop, car, anime, motion graphics, a long walk
and a reference-locked close-up — each one a different way to break a video model.
Four rules held it honest:
The prompt body is byte-identical across both models. Resolution, duration,
aspect and bitrate are API parameters, never prompt text. On the duration track the
RIG line reads model duration ceiling rather than a number, so even
there the two bodies match character for character.
One generation per cell. No re-rolls, no cherry-picking, no best-of-four.
Failures are published. The point of a bench is the breakage.
Anything a number cannot settle is marked OPEN and called on camera. You
cannot judge weight from a still frame, and this page will not pretend otherwise.
Schema is shot-builder-cinematic-action — a two-tier CARD + TAKE, every prompt
linted clean before it was sent.
02 — the results
Five findings
01 — the free flag
bitrate_mode: high costs nothing
Three to five times the data rate for zero extra credits, on both models. Preflighted on every call: the cost is identical to the credit. It also erased the worst artefact in the whole test. The default is the bug.
1.349 → 1.000
macroblocking on Seedance 2.0, same 36 credits
02 — the default tax
2.5 charges 44% more for less data
Five of six matched pairs had 2.5 shipping the thinner file while costing more. Finding 01 fixes this completely, and for free.
−34%
median bitrate vs 2.0, at 52 credits against 36
03 — the resolution trap
Seedance 2.0's 4K is starved
Nine times the pixels, 35% more data, 4.9× the price. Normalised back to 720p the 4K file carries less usable detail than the 1080p file. Buy 1080p.
0.025 bpp
against 0.167 at 720p — 6.6× fewer bits per pixel
04 — the grade test
Only the flat graphic broke
Contrast ×1.55, shadows to gamma 0.82, saturation ×1.35 — nothing posterised on the live-action tracks. Flat near-black fields are the exception: 2.0 blocked up badly until the free bitrate flag fixed it, and Gemini Omni Flash posterised the shadows worse than anything else in the test.
65 / 65
shadow levels held on every live-action clip after a hard grade
05 — the open question
2.5 moves half as much. Every time.
Either 2.5 is more controlled or it is doing less. A still frame cannot tell you which, so this one is left open and called on camera.
6 / 6
matched pairs where 2.5 had lower frame-to-frame delta
03 — the bench
Seven tracks
Drag the handle to wipe between the two models. Prompts and numbers are folded away under each one.
T1
IMPACT
Action · weight, contact, consequence
Tests what action prompts always fail at: mass, contact, and where the force goes. A punch that has to arrive after the weight that throws it.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
2.0 put 35% more detail on screen for 31% less money. On the numbers it takes this track outright.
Does the punch land with weight? Does the heel-plough throw spray that falls back inside frame, or does it float? Does either model twin the figures?
The prompt423 words · identical to both models
RIG: 8s · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
LOOK: 35mm anamorphic capture, 40mm, real grain and halation on the sodium, shallow falloff, imperfect focus. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Standing water breaks underfoot and throws spray that falls back inside the frame. Wet fabric clings and lags a beat behind the body. Nothing floats. Wardrobe, faces and container geography stay EXACT for the whole shot.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no crew or rig in the reflections.
OPEN: screen-left foreground, a broad man in a soaked canvas jacket already turning into the swing, back three-quarter to camera. Screen-right, mid-ground, a leaner figure in a wet grey hood is already stepping off the back foot. Water is still coming off both of them from the last exchange.
MOVE: one low push-in along the wet ground, a metre and no further.
RUN: the broad man's weight arrives first and the punch follows it → it catches the hooded figure high on the chest → the hooded figure folds over the force and goes back two stumbling steps → his heels plough standing water into a low sheet that hangs and drops → he stays up, one hand flat on a container wall → both stop, breathing, and the water keeps falling around them.
FACE: wants — to end it on this one
blocked by — the other man staying up
plays — jaw set, shoulders dropping as the swing spends itself, eyes never leaving him
turns on — the heels catching; he registers that it did NOT land clean
SOUND: <a dull heavy contact, more thud than crack> <boots tearing through standing water> <one hard exhale driven out on impact>
PIN: ONE impact and it lands on the chest. Both men stay in frame throughout. The spray falls back inside the frame. NEVER a third figure.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
2,933
3.11
107
1.065
1.051
13.37
52
2.0
1280x720
8.1
4,460
4.65
144
1.045
1.053
26.6
36
T2
SYNC
K-pop · four-body unison, the identity-drift cliff
Deliberately walks into the documented failure mode: three or more simultaneous characters. Four dancers hitting the same beat is the hardest ask in the test.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
The only pair where 2.5 shipped the fatter file, and it led on detail too. 2.5's 30-image reference ceiling against 2.0's 9 should matter most exactly here.
Count the dancers every second — does either produce a fifth? Do all four hit the beat together? Do the four faces stay distinct or converge?
The prompt428 words · identical to both models
RIG: 8s · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
LOOK: 35mm anamorphic capture, 40mm, real grain and halation on the sodium, shallow falloff, imperfect focus. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Standing water breaks underfoot and throws spray that falls back inside the frame. Wet fabric clings and lags a beat behind the body. Nothing floats. Wardrobe, faces and container geography stay EXACT for the whole shot.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no crew or rig in the reflections.
OPEN: four dancers in a shallow V facing camera, already mid-count — arms up, wrists crossed, weight loaded on the back foot. Matching black technical wear, soaked at the shins. The lead is centre-foreground; the other three sit half a step behind, screen-left, screen-right and deep centre.
MOVE: one slow dolly back, staying square to the V.
RUN: all four drop the crossed arms to the hip on the same beat → they pivot as one a quarter turn to screen-left and hold → the lead breaks forward a single step alone while the other three drop to a crouch → the lead snaps a head turn back to camera → all four rise into the shape they opened on and hold it.
FACE: wants — to be read as one body and not as four people
blocked by — the wet ground stealing grip on every pivot
plays — chins level, eyes locked to lens, controlled breath showing in the cold
turns on — the lead's break forward; the three behind become a wall
SOUND: (a hard four-on-the-floor beat, sub-heavy, no melody) <four pairs of shoes striking wet ground on the SAME frame> <fabric snapping on the turn>
PIN: all four land EVERY beat together except the lead's single break. Four dancers and NEVER a fifth. The four faces stay distinct from one another. The V never flattens into a line.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
4,715
4.9
169
1.040
1.058
13.15
52
2.0
1280x720
8.1
4,279
4.46
151
1.042
1.050
24.94
36
T3
LACQUER
Car commercial · specular surface and reflection
No people. Wet black paint under sodium is the most demanding surface in commercial work: one continuous highlight travelling a flank, and a reflection welded to the car. This track carried the bitrate and resolution ladder.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s6 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
Six renders. The bitrate flag multiplied the data rate 5× at zero cost; the 4K option cost 4.9× and delivered 6.6× fewer bits per pixel than 720p.
Does the highlight travel as one unbroken run down the flank, or break and re-form wrongly at the shut-line? Does the reflection stay locked to the body? Any invented badge or plate despite the ban?
The prompt392 words · identical to both models
RIG: 8s · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
LOOK: 35mm anamorphic capture, 40mm, real grain and halation on the sodium, shallow falloff, imperfect focus. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Standing water breaks underfoot and throws spray that falls back inside the frame. Wet fabric clings and lags a beat behind the body. Nothing floats. Wardrobe, faces and container geography stay EXACT for the whole shot.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no crew or rig in the reflections.
OPEN: a low wide dark saloon side-on across the frame, nose to screen-right, already rolling forward at walking pace. Camera sits at hub height a metre off the wet ground, so the car and its full reflection fill the frame together.
MOVE: one lateral track alongside the car, matched to its speed, holding the same point on the flank.
RUN: the sodium floods travel down the flank as one continuous highlight → the highlight breaks at the door shut-line and re-forms past it → the front wheel throws a fine sheet of water outward that catches the blue lamp → the car clears the last container and the sodium falls off the body → it settles into the cold blue and comes to rest, and the reflection settles a beat after the car does.
FACE: no people in this shot — the performance is the surface.
SOUND: <tyres peeling slowly off wet asphalt> <a low engine note held steady, no revs> <water sheeting off the wheel arch>
PIN: ONE unbroken highlight travels the flank. The reflection stays locked to the car. No badge, no logo, no number plate. NEVER a driver, NEVER a person in frame.
The numbers6 renders · bitrate, detail, macroblocking
2.5 · high2.0 · high2.0 · 1080p2.0 · 4K
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
2,224
2.4
193
1.079
1.083
5.85
52
2.0
1280x720
8.1
3,702
3.88
90
1.073
1.114
26.91
36
2.5 · high
1280x720
8
11,301
11.52
102
1.253
1.314
23.89
52
2.0 · high
1280x720
8.1
10,398
10.61
165
1.080
1.099
29.23
36
2.0 · 1080p
1920x1080
8.1
7,595
7.8
147
1.039
1.055
35.51
72
2.0 · 4K
3840x2160
8
5,035
5.22
118
1.004
1.047
19.23
176
T4
INK
2D anime · non-photoreal style hold
The same fight as IMPACT, re-lit as hand-drawn cel animation on twos. Tests whether a model can hold a non-photoreal style for eight seconds without drifting back toward photography.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
2.5 put 94% more edge energy on screen at the same data rate. For cel animation, edge energy is ink line. The largest single-metric gap anywhere in the test.
Is 2.5's extra line energy clean ink, or is it boiling — line-weight jitter frame to frame? Does either slip back toward photoreal?
The prompt436 words · identical to both models
RIG: 8s · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
LOOK: hand-drawn 2D anime cel animation on twos. Ink outline of varying weight, flat cel shading in exactly two tones per surface, hard-edged shadow shapes, painted background with visible brush texture. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Water is drawn as shaped highlight cels, NEVER photographic spray, and falls back inside the frame. Wet fabric clings and lags the body. Nothing floats. Line weight, character design and geography stay EXACT throughout.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no photoreal frames, no 3D render look.
OPEN: screen-left foreground, a broad man in a soaked canvas jacket already turning into the swing, back three-quarter to camera. Screen-right, mid-ground, a leaner figure in a wet grey hood is already stepping off the back foot. Water is still coming off both of them from the last exchange.
MOVE: one low push-in along the wet ground, a metre and no further.
RUN: the broad man's weight arrives first and the punch follows it → it catches the hooded figure high on the chest → the hooded figure folds over the force and goes back two stumbling steps → his heels plough standing water into a low sheet that hangs and drops → he stays up, one hand flat on a container wall → both stop, breathing, and the water keeps falling around them.
FACE: wants — to end it on this one
blocked by — the other man staying up
plays — jaw set, shoulders dropping as the swing spends itself, eyes never leaving him
turns on — the heels catching; he registers that it did NOT land clean
SOUND: <a dull heavy contact, more thud than crack> <boots tearing through standing water> <one hard exhale driven out on impact>
PIN: ONE impact and it lands on the chest. Line weight stays EXACT on both characters. The spray is drawn, never photographic. NEVER a third figure.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8
3,597
3.78
511
1.012
1.030
21.84
52
2.0
1280x720
8.1
3,718
3.9
263
1.020
1.029
31.02
36
T5
TITLE
Motion graphics · on-screen text, the known weak spot
Written in motion-builder grammar rather than the live-action CARD: style lock, action-only shot block, quoted text strings, sound design only. On-screen text is the weakest area of every Seedance version.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s6 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question · caveat
Six renders, four engines. At default bitrate 2.0 is the worst file in the episode — 1.349 macroblocking, rising to 1.784 under grade. With the free high flag it becomes the best: 1.000, no detectable block structure, and detail nearly doubled. Gemini Omni Flash is by far the cheapest at 24 credits and holds blocking better than 2.0-default, but posterises the near-black worst of anything here — 48 distinct shadow levels against 57-65 everywhere else.
The real question is legibility. Do the four quoted strings render as letterforms or as corrupted glyphs? Rate each 0–10. If both fail, that is the finding — generate clean, set type in post.
Four engines on the identical prompt: Seedance 2.5, Seedance 2.0, Google Gemini Omni Flash and Higgsfield Cinema Studio 3.0 — all run through Higgsfield, all 8s 720p 16:9.
The prompt205 words · identical to both models
STYLE LOCK: dark dock-side title card system. Near-black ground with visible film grain. One sodium-orange accent and one cold blue accent, nothing else. Type is heavy condensed uppercase sans, tightly tracked, always locked to a hard left margin. Rules are one-pixel hairlines. Surfaces are flat matte, no gloss, no bevel, no drop shadow. Motion is mechanical and abrupt — cuts and snaps, never eases. Do NOT copy any reference layout; take only the visual language.
SHOT: on a near-black field, a single hairline rule draws left to right across the lower third and stops → the words snap on one word at a time from the left margin, each landing on a hard beat with no fade → a solid orange block wipes in from screen-right behind the last word and holds → the hairline drops away → the block collapses to a hairline and the words cut out together.
TEXT ON SCREEN: "SEEDANCE" (H1), "2.5" (H1), "VS" (label), "2.0" (H1)
SOUND: <a mechanical shutter tick on each word landing> <one low sub drop when the block wipes in> <tape hiss underneath, constant>
AVOID: no music, no voice, no gradients, no glow, no lens flare, no 3D extrusion, no letters other than the quoted strings.
The numbers6 renders · bitrate, detail, macroblocking
2.5 · high2.0 · highOmni FlashCinema Studio 3.0
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
645
0.81
137
1.096
0.870
5.34
52
2.0
1280x720
8.1
2,632
2.81
161
1.349
1.784
6.25
36
2.5 · high
1280x720
8.1
4,601
4.79
112
1.015
1.188
7.2
52
2.0 · high
1280x720
8.1
7,018
7.21
306
1.000
0.740
4.38
36
Omni Flash
1280x720
8
2,035
2.18
81
1.134
0.947
8.28
24
Cinema Studio 3.0
1280x720
8.1
1,587
1.74
80
1.319
1.160
3.12
40
T6
LONG
Duration ceiling · 20s continuous against a 15s wall
2.5 tops out at 30 seconds, 2.0 at 15. Both got the same continuous follow down the container corridor and were asked for their maximum.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 130 cr2.0 · 68 cr1280x720 · 20.1s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
2.5 sustained five extra seconds and doubled the edge energy, for roughly double the price. Past fifteen seconds, 2.0 cannot compete at any price.
Where does each start to drift? Mark the timecode. Does the hood stay up the whole way? Does 2.5's last third hold together?
The prompt420 words · identical to both models
RIG: model duration ceiling · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
LOOK: 35mm anamorphic capture, 40mm, real grain and halation on the sodium, shallow falloff, imperfect focus. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Standing water breaks underfoot and throws spray that falls back inside the frame. Wet fabric clings and lags a beat behind the body. Nothing floats. Wardrobe, faces and container geography stay EXACT for the whole shot.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no crew or rig in the reflections.
OPEN: a lone figure in a wet grey hood walking away from camera down the centre of the container corridor, already three strides in, hands loose at the sides. Camera is behind at shoulder height.
MOVE: one continuous follow behind the figure, holding the same distance the whole way.
RUN: the figure walks the corridor and each sodium flood passes over the shoulders in turn → at the third container he steps around a standing pool without breaking stride → a gull lifts off a container top screen-right and he does NOT look at it → the corridor opens onto black water and the sodium runs out → he stops at the edge with his back still to camera → he turns his head a quarter to screen-left and holds there.
FACE: wants — to reach the water without being seen
blocked by — the length of the corridor and the light he has to cross
plays — steady pace, shoulders low, head still, hands never coming up
turns on — the sodium running out; the shoulders finally drop
SOUND: <boots on wet asphalt, a steady unhurried rhythm> <one gull, distant, once> <water lapping steel at the corridor mouth>
PIN: ONE unbroken take, NEVER a cut. The hood stays up EVERY frame. The face is NEVER shown. The walk pace never changes.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
20.1
3,141
8.23
120
1.103
1.082
22.9
130
2.0
1280x720
15.1
3,492
6.85
58
1.109
1.102
22.72
68
T7
LOCK
Reference identity · 30-image net against 9
A three-view character sheet was generated first, then bound into both models with explicit exclusions — NOT the pose, NOT the framing, NOT the studio light, NOT the backing card. 2.5 ran omni_reference, 2.0 image_references.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s2 renders
Image reference 1 reference · what went in
A three-view character sheet was generated first (Nano Banana Pro — hood up, scar through the right eyebrow), then bound into both models with an explicit exclusion clause: NOT the pose, NOT the framing, NOT the flat studio light, NOT the grey backing card. 2.5 ran in omni_reference mode, 2.0 with image_references.
Character sheet on a flat neutral mid-grey card, even studio lighting, three views left to right:
full-body front, full-body left profile, tight head-and-shoulders with the hood UP. The SAME man
at the SAME scale in all three. Male, late twenties, lean and wiry, 178cm, medium-brown skin,
short dark hair, a small vertical scar breaking the right eyebrow. Wears a soaked heather-grey
cotton hoodie with the hood up and the rim heavy with water, a black long-sleeve beneath, dark
tapered work trousers, black boots worn at the toe. Neutral A-pose in the two full-body views.
Sharp focus throughout, no motion blur, no colour grade, no props, no text or labels anywhere.
What the numbers say
verdict · open question
Low detail on both is expected: a locked-off close-up with shallow falloff has little edge energy by design. The identity question is not a number.
Pure eyeball. Put the reference sheet next to both. Is the eyebrow scar there? Same face? Did either import the flat studio light or the grey backing card despite the exclusions?
The prompt455 words · identical to both models
RIG: 8s · 720p · 16:9 · audio on · 24fps
SET: a container dock after rain, past midnight. Stacked steel containers make a corridor two lanes wide, the far end open to black water. Standing water sheets the asphalt and holds every light twice. The ONLY sources are four sodium floods on the container tops throwing hard orange from high and behind, and one cold blue-white work lamp on a tripod at the near end, low and to camera-left.
CAST: @Image1 → the hooded man: face, skin, the scar through the right eyebrow, the grey hoodie and its hood rim. NOT the pose, NOT the framing, NOT the flat studio light, NOT the grey backing card, NOT the three-panel layout.
LOOK: 35mm anamorphic capture, 40mm, real grain and halation on the sodium, shallow falloff, imperfect focus. Palette 60 wet black asphalt and steel, 25 sodium orange, 15 cold blue.
LAW: real gravity and inertia — weight arrives before the limb that carries it, and every force lands somewhere. Standing water breaks underfoot and throws spray that falls back inside the frame. Wet fabric clings and lags a beat behind the body. Nothing floats. Wardrobe, faces and container geography stay EXACT for the whole shot.
BED: <rain runoff ticking off steel> <a distant harbour horn> <the work lamp ballast humming>
BAN: no extra people, no cuts, no text or captions, no crew or rig in the reflections.
OPEN: the hooded figure alone, medium-close, three-quarter to camera at screen-right, already breathing hard from a fight that has just stopped. The blue work lamp is behind him, so the hood rim is edged cold and the face sits in sodium spill.
MOVE: locked off. The frame does NOT travel — only the operator's breathing shows.
RUN: he lifts his head and the hood rim clears his eyes → he looks off screen-left at whatever is still standing there → he wipes the back of one wrist across his mouth and checks it → he sets his jaw and squares his shoulders back to where he was looking → he holds, and only the breath keeps moving.
FACE: wants — to look like he can go again
blocked by — a body that has already decided otherwise
plays — nostrils working, one slow blink, a swallow he tries to hide
turns on — the look at his own wrist; the bravado drops for one frame and comes straight back
SOUND: <breath through the nose, uneven, close to the mic> <water dripping off the hood onto the shoulder> <one wet footstep from somewhere off-frame>
PIN: the SAME face and the SAME hood as the reference in EVERY frame. Camera never moves. He is alone — NEVER a second figure.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
1,895
2.07
37
1.090
1.094
9.81
52
2.0
1280x720
8.1
2,989
3.17
44
1.034
1.061
17.83
36
03b — complexity
Advanced
Maximum segments, dynamic camera, saturated looks and real production grammar. Six segments is inside 2.5's 5–6 shot budget and past 2.0's 3–4 — the one place the two models are specified to diverge.
A1
ARCANE
Painterly action · six segments, cut-budget ceiling
Hand-painted moving oil painting on a cobalt snow-steppe. Six segments, five hard cuts asked for — inside 2.5's 5-6 shot budget and deliberately past 2.0's 3-4. This is the one place the two models are specified to diverge, and the simple tier could not test it.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8.1s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
Both models cut far MORE than asked: 8 detected discontinuities against 5 requested, on both. A dense segmented prompt does not just permit cutting, it provokes it. Neither model respected the stated count.
Watch whether the extra cuts are real edits or the scene detector firing on the whip pan and the ember flicker. Does the painterly look actually hold on twos, or does it smooth into 3D? Does the wolf stay one wolf?
The prompt1042 words · identical to both models
STYLE LOCK
Hand-painted 2D animation in the Arcane tradition — a moving oil painting, NOT cel-shaded anime,
NOT 3D, NOT photoreal, NO CGI smoothness, NO photobash. Every frame carries visible brushwork:
loose painterly strokes on the backgrounds, thick impasto on the light, textured canvas tooth
throughout. Characters are painted with soft-edged rendering and a warm sculpted key, held
together by a dark broken contour line rather than a clean ink outline. Rich saturated palette:
deep cobalt and violet snow-dusk, ember orange and hot gold firelight, blood-crimson accents.
Strong chiaroscuro — deep shadow with colour inside it, never flat black. Animated at 12fps ON
TWOS with constant painterly boil; atmospheric elements (snow, breath vapour, embers, smoke)
move SMOOTHLY while figures step on twos. NO smooth interpolation, NO motion blur, NO morphing,
NO liquid or melting surfaces.
CONTINUITYTHE RIDER: a woman in her thirties, weathered brown skin, sharp cheekbones, black hair in a
thick braid whipped loose at the temples, a white cold-scar across the bridge of her nose. Heavy
felted wool coat in oxblood red with a wide fur collar rimed white with frost, leather bracers,
a dark fur hat. THE WOLF: a huge grey steppe she-wolf, silver-backed with a black saddle mark,
one ear notched, amber eyes. Both designs, their palette and their brush treatment stay identical
across every segment. SETTING: a high winter steppe at blue hour — a wide snowfield running to
black mountains, a nomad camp of three round felt yurts on the ridge behind, each glowing amber
from within, a bonfire throwing embers up into falling snow. Geography and light positions stay
fixed across cuts.
DIRECTION
One camera action per segment, stated explicitly. All cuts are HARD CUTS on a named beat. Action
is described by mass and consequence, never by speed. Sound cues per segment. No music. Only SFX.
SHOT — six segments, five hard cuts, 8 seconds
SEGMENT 1 — WIDE ESTABLISH (0.0s–1.4s)
Painted wide of the snowfield at blue hour, the three glowing yurts small on the ridge, the
bonfire throwing embers. THE RIDER stands alone mid-field in the oxblood coat, back to camera,
facing the black mountains. Snow falls smoothly through a frame that is otherwise stepping on twos.
CAMERA: locked off, the frame does NOT travel.
SFX: <wind moving low across open snow> <fire crackling small and distant>
HARD CUT.SEGMENT 2 — THE WOLF ARRIVES (1.4s–2.8s)
Low painted wide from snow level, the drifts filling the bottom third. THE WOLF comes over a rise
at a heavy lope, silver back catching the last cobalt light, breath vapour streaming smooth behind
a body that steps on twos. Snow bursts off her paws as painted flat shapes.
CAMERA: a hard whip pan tracking with her, ending as she plants.
SFX: <deep paws compressing dry snow> <one low chest growl>
HARD CUT on the plant.
SEGMENT 3 — THE RIDER TURNS (2.8s–4.0s)
Painted medium on THE RIDER, ember-orange firelight raking her from screen-right, cobalt fill from
the sky on the other side, the fur collar rimed white and edge-lit hot.
CAMERA: locked off, a slow painterly push in, ending on her eyes.
ACTION: she turns her head over her shoulder toward the wolf, the braid swinging late and landing
two frames after the head stops.
SFX: <heavy wool and frost shifting> <a caught breath>
HARD CUT.SEGMENT 4 — EYE MATCH, THE WOLF (4.0s–5.0s)
Painted close on THE WOLF's face, amber eyes catching the fire, the notched ear, individual guard
hairs rimed with frost picked out in thick paint.
CAMERA: locked off. The frame does NOT travel.
ACTION: her head lowers a fraction and the amber eyes hold, unblinking. Breath vapour rolls
smoothly from her muzzle.
SFX: <a single low breath through a wet muzzle>
HARD CUT.SEGMENT 5 — THEY CLOSE (5.0s–6.6s)
Painted two-shot, profile, both of them in frame with the bonfire burning between them and camera
so embers cross the foreground as smooth-moving painted sparks.
CAMERA: an arc around them at chest height, staying level, ending square to both.
ACTION: THE RIDER takes three deliberate steps toward the wolf through knee-deep snow, weight
sinking and lifting on each one. The wolf holds absolutely still and lets her come.
SFX: <snow compressing under boots, three times> <embers popping close to lens>
HARD CUT.SEGMENT 6 — THE HAND (6.6s–8.0s)
Painted close on the space between them: the rider's bare hand entering frame from screen-left,
the wolf's muzzle from screen-right, the gap between them lit hot gold by the fire.
CAMERA: locked off. The frame does NOT travel.
ACTION: the hand stops a palm's width short and stays there. The wolf leans in and closes the last
of the gap herself. Snow falls smoothly through the hot gold light between them.
SFX: <wind dropping away to almost nothing> <one long exhale>
PHYSICS
Real weight and real cold, expressed in paint. Snow compresses and holds the shape of what pressed
it. Deep snow drags at every step and the body lifts to clear it. Heavy felted wool swings late and
holds its folds. The braid and the fur collar carry momentum and settle two frames after the body.
Breath vapour and embers rise and drift smoothly on the air while every figure steps on twos.
Nothing floats, nothing glides, nothing resets between cuts.
LIGHTING
Two sources only: the bonfire screen-right, low and warm, throwing ember-orange key and live
flicker across faces and fur; and the last cobalt blue-hour sky above, filling every shadow with
cold colour. The three yurts glow amber from within on the far ridge as small practical accents.
Shadows stay deep but always carry colour, never flat black. Faces stay readable at all times.
POSITIVE LOCKS
Hand-painted moving-oil-painting look in every single frame, animated on twos with visible
painterly boil and canvas texture. Exactly six segments and exactly five hard cuts. The SAME rider
design and the SAME wolf design in every segment they appear, matching the CONTINUITY block in
face, wardrobe, markings and palette. Exactly one woman and exactly one wolf for the whole clip.
Snow and embers move smoothly while figures step on twos. The bonfire stays screen-right and the
yurts stay on the far ridge across every cut. The final frame is the hand and the muzzle touching.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8.1
10,570
10.78
0
0.000
0.000
0
52
2.0
1280x720
8.1
10,335
10.55
0
0.000
0.000
0
36
A2
IDOL
K-pop performance · five bodies, saturated stage, beat-locked cuts
Mirror-black stage under a magenta and cyan LED wall, five dancers in white and chrome. Five segments, four cuts asked for, every cut on a kick drum. Five bodies is well past the three-character identity-drift cliff.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 52 cr2.0 · 36 cr1280x720 · 8s2 renders
Text to video no input assets · the prompt is the whole input
No input assets. The prompt below is the entire input — copy it and you can reproduce this render from nothing.
What the numbers say
verdict · open question
The most cut-heavy pair in the test: 10 detected on 2.5 and 14 on 2.0, against 4 requested. 2.0 cuts more aggressively than 2.5 on every single advanced track.
Count the dancers per segment — five, or does a sixth appear? Do the cuts actually land on the kick? Is the whip pan a real pan or a cut disguised as one?
The prompt1248 words · identical to both models
SCENE CONTEXT
Live-action photoreal K-pop performance video, 8 seconds, a single 8-count performed by five
dancers on a mirror-black stage in front of a floor-to-ceiling LED wall. It opens on a locked
symmetrical wide, then cuts on the kick to a low hero angle, a whip-pan across the line, a locked
beauty close-up, and a crane-back finish. Every cut lands exactly on a kick drum. The mood is
expensive, saturated and precise — a title-track performance take, not a rehearsal.
ACTIVE REFERENCESTHE LEAD: female dancer, early twenties, 168cm, athletic and compact, East Asian, sharp jawline,
black hair in a high tight ponytail with a blunt fringe, a single silver hoop in the left ear.
She wears a structured white cropped moto jacket with oversized silver hardware over a black mesh
top, high-waisted white technical trousers with a chrome chain at the right hip, and white
low-profile boots. She is centre of the formation in every wide.
DANCERS 2–5: four more dancers in the SAME white-and-chrome outfit, each a visibly DIFFERENT
person from the lead and from each other — DANCER 2 taller with a platinum buzz cut, DANCER 3
shorter with a long black high ponytail, DANCER 4 mid-height with a chin-length copper bob,
DANCER 5 tall with black hair in twin braids. Five dancers in total and never a sixth.
THE STAGE: a seamless mirror-black floor that reflects everything above it at full strength, and
behind the dancers a floor-to-ceiling LED wall running a hard geometric pattern of magenta and
cyan bars. Six moving-head spotlights above throw hard white beams down through a light haze.
LOCATION MAP (exact positions)
The five dancers stand in a straight line across frame, evenly spaced, all facing camera. THE LEAD
is centre. DANCER 2 and DANCER 4 are to her screen-left, DANCER 3 and DANCER 5 to her screen-right.
The LED wall fills the entire background behind them. The mirror floor carries a full inverted
reflection of all five and of the LED wall. Haze in the air makes every spotlight beam visible as
a solid shaft.
GAZE AND EYELINES
All five look directly down the lens on every held beat — chins level, unblinking, performance
direct. Eyes leave the lens only during the turn in Segment 3 and snap back on the landing. No
dancer ever looks at another dancer.
FORMAT MODE
Controlled five-segment sequence with HARD CUTS, every cut landing on a kick drum. Real-time
motion throughout except Segment 4, which is a controlled ramp. No dialogue.
SEGMENT 1 — LOCKED SYMMETRICAL WIDE (0.0s–1.8s)
Full-body wide, camera locked and dead centre, all five in frame head to boot with their full
reflections in the mirror floor and the magenta-cyan LED wall behind.
ACTION: on the first kick all five snap from a low crouch up into a hard arms-crossed pose, landing
on the SAME frame. The LED wall flashes to full white for exactly two frames on the hit and returns
to magenta and cyan.
LENS: locked widescreen wide, deep focus, clean crisp glass with a matte finish, spotlight beams
reading as solid haze shafts.
HARD CUT on the kick.
SEGMENT 2 — LOW HERO ANGLE (1.8s–3.4s)
Camera at floor level looking up, the mirror floor in the near foreground doubling the line of
bodies, the spotlight beams converging overhead.
ACTION: all five drive one arm down and across on the beat and drop into a wide second-position
stance. Chrome hardware catches the moving heads and throws hard specular glints.
LENS: a slow dolly in along the floor, ending as the stance locks; widescreen, low, clean crisp
glass with a matte finish.
HARD CUT on the kick.
SEGMENT 3 — WHIP PAN ACROSS THE LINE (3.4s–4.8s)
Chest height, square to the line.
ACTION: all five execute a full body turn in unison, and the camera whips screen-left to
screen-right across the whole line, arriving on the lead exactly as she completes the turn and
faces front. Ponytails and braids carry the rotation and land late.
LENS: a hard whip pan, motion-smeared at the midpoint and resolving sharp on the lead.
HARD CUT on the kick.
SEGMENT 4 — LOCKED BEAUTY CLOSE-UP, THE LEAD (4.8s–6.4s)
Locked facial close-up on the lead, magenta rimming one side of her face and cyan the other, a
white spotlight raking her cheekbone from above.
ACTION: she holds the lens dead-on, then delivers one sharp isolated head tilt. The motion runs at
half speed for the first half of the segment and snaps back to real time on the landing.
LENS: locked off, the frame does NOT travel; widescreen close-up, very shallow, the LED wall
dissolved to soft magenta and cyan bokeh behind her.
HARD CUT on the kick.
SEGMENT 5 — CRANE BACK, THE FINISH (6.4s–8.0s)
Opening on the lead and widening to reveal all five again.
ACTION: all five hit the final pose together — one arm up, chin down — and hold it dead still. The
LED wall resolves to a single hard magenta field behind them. The haze settles through the beams.
LENS: a crane up and back off the lead to a full symmetrical wide, decelerating and stopping
exactly as they hit the pose.
PHYSICS
Real weight and real momentum. Every landing is absorbed through the knees. Ponytails, braids,
jacket hems and the chrome hip chain carry momentum and settle a beat after the bodies that swing
them. Structured jacket fabric holds its shape and does not cling. The mirror floor reflection
tracks every dancer exactly. Haze moves in real air currents disturbed by the bodies passing
through it. Nothing floats, nothing glides, nothing resets between cuts.
LIGHTING
Six overhead moving-head spotlights throwing hard white beams down through visible haze, plus the
full-height LED wall behind running magenta and cyan as a saturated backlight that rims every
dancer. The mirror floor bounces both back up under their chins. High contrast, deeply saturated,
faces always readable. Every light source renders as a soft, round, contained glow within its own
source.
AUDIO
A hard four-on-the-floor beat with a heavy sub and a bright synth stab, no vocals, running unbroken
through all five segments, with every hard cut landing exactly on a kick drum. Under it, diegetic
only: five pairs of boots striking a hard mirror floor on the SAME frame, structured fabric
snapping on each turn, chrome hardware ticking, and controlled breathing.
STYLE
Fully photoreal live-action music video, 35mm filmic capture, widescreen 2.39:1 frame, clean crisp
lens glass with a matte finish, natural shallow depth of field on the close-ups, fine organic
grain, richly saturated high-contrast grade with magenta and cyan dominant and clean white
highlights, premium title-track polish. The image stays sharp, steady and clean.
POSITIVE LOCKS
Exactly five dancers for the entire clip, in every segment, and never a sixth. All five wear the
identical white-and-chrome outfit and all five remain visibly different people from one another in
height, build and hair across every cut. The lead matches her reference description in face, hair
and wardrobe in every frame she appears. Exactly five segments and exactly four hard cuts, each cut
landing on a kick drum. All five land every beat on the same frame in Segments 1, 2, 3 and 5. The
line stays a straight evenly-spaced line in every wide. The mirror floor carries a full reflection
of all five in every segment it appears. The LED wall stays behind the dancers across every cut.
All five hold the final pose dead still as the clip ends.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1280x720
8
12,207
12.43
0
0.000
0.000
0
52
2.0
1280x720
8.1
12,705
12.93
0
0.000
0.000
0
36
A5
GRAVITY
Full 15s commercial · reference-bound brand work, sound design only
A complete 15-second Mercedes-Benz spot, end to end, in an abstract void. Six segments: macro emblem tease, orbit reveal, rigged wheel, overhead colour field, headlight signature, pack shot. The badge comes from a bound reference plate, never from the prompt text — and the final beat is left as a clean plate for the wordmark to be composited in post.
15s is 2.0's absolute ceiling and mid-range for 2.5. 2.5 shipped the thinner file again (7,845 kbps against 9,863) — the same default-bitrate pattern as the simple tier, holding even at high bitrate and 15 seconds.
The brand test: is the three-pointed star intact and correctly formed in every frame it appears, on both? Does the car stay the same car across six cuts? Is the pack shot actually static and usable as a plate?
The prompt1485 words · identical to both models
SCENE CONTEXT
Hyper-real automotive commercial, 15 seconds, end to end, no people anywhere. A single production
Mercedes-Benz coupé in an abstract void studio — no road, no city, no landscape, only light,
colour and reflection. The spot opens on a macro tease of the emblem, reveals the whole car in an
orbit as the world ignites around it, drives through a field of pure colour, resolves on
individual light signatures, and lands on a static hero pack shot. The mood is premium, confident
and cinematic — a broadcast title spot, not a dealer film.
ACTIVE REFERENCES
@veh_MB_hero_v1: a low wide modern Mercedes-Benz coupé in liquid graphite-grey metallic paint with
a mirror-deep clearcoat. Large upright grille carrying the three-pointed star emblem centred on it,
slim horizontal LED headlights with three internal light bars each, long bonnet with two subtle
power domes, flush door handles, deep-dish multi-spoke forged wheels in gloss black with polished
rims, low-profile tyres, a taut shoulder line running the full flank. The SAME car, the same paint,
the same wheels and the same emblem in every segment. 100% matches the reference.
THE VOID: a seamless glossy black infinity floor with no horizon line and no walls, filled with
fine atmospheric haze. Colour exists only as light: a deep amber-orange gradient wash from the
lower left and a cold electric-cyan rim source from the rear right. One enormous soft overhead
strip light runs the length of the space. The floor carries a full mirror reflection of the car
and of every colour source at all times.
LOCATION MAP
The car sits at the exact centre of the void on the glossy black floor. The amber wash enters from
lower screen-left, the cyan rim from upper screen-right behind the car, and the overhead strip
light runs front to back directly above the car's centre line. The mirror floor doubles everything.
Haze makes both coloured sources read as soft volumetric glows.
FORMAT MODE
Controlled six-segment commercial sequence with HARD CUTS. Real-time motion except Segment 4, a
controlled ramp. No people, no dialogue, no on-screen text. Sound design only.
SEGMENT 1 — MACRO TEASE, THE EMBLEM (0.0s–2.2s)
Extreme macro on the grille, so tight only a fragment of polished metal is readable against black.
ACTION: a hard specular highlight travels slowly across the chrome, and the three-pointed star
resolves out of darkness as the light crosses it. The rest of the frame stays black.
LENS: locked macro with a slow rack focus pulling the emblem from soft to razor sharp; extremely
shallow, clean crisp glass with a matte finish.
SFX: <a deep sub-bass swell rising from silence> <a single metallic resonance as the light crosses>
HARD CUT on the star reaching full sharpness.
SEGMENT 2 — THE REVEAL, ORBIT (2.2s–5.4s)
The full car, three-quarter front, alone in the black void.
ACTION: as the camera moves, the amber wash ignites from lower screen-left and the cyan rim fires
from behind screen-right, and the whole void resolves from black into colour around the stationary
car. The overhead strip light lays one unbroken white highlight down the bonnet and roof. The
mirror floor fills with both gradients.
LENS: a smooth orbit around the car at headlight height, staying level, decelerating and ending
square to the three-quarter front; widescreen, clean crisp glass with a matte finish.
SFX: <the sub-bass resolving into a low sustained tone> <a soft ignition swell as the light arrives>
HARD CUT.SEGMENT 3 — RIGGED LOW, THE WHEEL (5.4s–7.8s)
Camera rigged to the car itself at the front arch, looking back along the flank at wheel height.
ACTION: the car is now moving. The wheel turns under load, the polished rim strobing, and the
glossy black floor beneath streams past as a river of liquid amber and cyan reflection. The
shoulder line holds one continuous travelling highlight.
LENS: locked to the body, no independent camera movement; widescreen, shallow, clean crisp glass.
SFX: <a low mechanical whirr under the tone> <air moving fast and close>
HARD CUT.SEGMENT 4 — OVERHEAD, THE COLOUR FIELD (7.8s–10.4s)
Directly overhead, top-down, the car small and centred, travelling through a field of pure
saturated colour — broad vertical bands of deep crimson, cobalt and burnt orange streaming past
beneath and around it as pure abstraction, no road markings, no geography.
ACTION: the car cuts a clean line through the colour field. The motion runs at half speed for the
first half of the segment and snaps back to real time as the car crosses centre frame.
LENS: tracking directly above the car, matched to its speed; widescreen, deep focus.
SFX: <the tone dropping to a filtered rumble on the ramp and snapping back on the return>
HARD CUT.SEGMENT 5 — THE LIGHT SIGNATURE (10.4s–12.6s)
Locked tight and low on the front of the car, three-quarter, the headlight filling most of frame.
ACTION: the three internal light bars of the headlight ignite one after another, left to right,
each one throwing a hard cyan spill across the glossy floor. The grille and the emblem sit just
inside frame edge, catching amber.
LENS: locked off, the frame does NOT travel; widescreen, very shallow, clean crisp glass.
SFX: <three soft electrical ignitions in sequence> <the sustained tone rising underneath>
HARD CUT.SEGMENT 6 — PACK SHOT (12.6s–15.0s)
The hero frame. Full car, three-quarter front, dead centre-left of frame with clean empty negative
space in the upper right of the void.
ACTION: the car comes smoothly to rest. The suspension settles a beat after it stops. The
reflection in the mirror floor settles a beat after the suspension. The haze drifts and stills.
Everything holds, absolutely static, for the last full second.
LENS: a slow crane up and back to the hero three-quarter, decelerating and stopping exactly as the
car settles; widescreen, deep focus, clean crisp glass with a matte finish.
SFX: <everything resolving to one clean sustained low tone> <a single soft impact as it settles>
<near silence for the final held second>
PHYSICS
Real mass and real suspension. The car settles on its springs when it stops and the nose dips
fractionally. Tyres deform very slightly at the contact patch under load. The mirror-floor
reflection is a true inverted mirror of the car and tracks it exactly, distorting only where the
floor is disturbed. Haze parts around the car in motion and closes behind it. Light behaves
physically: specular highlights travel across curved panels as one continuous form, and every
reflection in the paint is a real reflection of a real source in the void.
LIGHTING
Three sources only. One enormous soft overhead strip light running front to back above the car,
laying a single unbroken white highlight down the bonnet and roof. A deep amber-orange gradient
wash filling the void from lower screen-left. A cold electric-cyan rim source from upper
screen-right behind the car, separating the roofline and rear haunch from the black. The glossy
black floor mirrors all three. High dynamic range with deep blacks that still hold detail. Every
light source renders as a soft, contained glow within its own source.
AUDIO
Sound design only — no music, no voice, no dialogue. A deep sub-bass swell rising from silence and
resolving into one low sustained tone that carries the whole spot. Under it: a single metallic
resonance on the emblem, a soft ignition swell on the reveal, a low mechanical whirr and fast close
air in motion, three soft electrical ignitions in sequence on the headlight, one soft impact as the
car settles, and near silence for the final held second.
STYLE
Fully photoreal hyper-real automotive commercial, large-format digital cinema capture, widescreen
2.39:1, clean crisp lens glass with a matte finish, natural shallow depth of field on the macro and
detail beats, deep focus on the wides, fine organic grain, high dynamic range, deeply saturated
amber-and-cyan grade against true black, premium broadcast campaign finish. Every light source
renders as a soft, contained glow within its own source. The image stays sharp, steady and clean.
POSITIVE LOCKS
The SAME car in every segment — one body, one graphite-grey paint, one wheel design, one emblem —
matching @veh_MB_hero_v1 exactly. The three-pointed star emblem appears exactly as it does in the
reference and stays crisp and correctly formed wherever it is visible. The frame contains the car,
the void, the light and the floor reflection only, for the entire clip — no people, no road, no
buildings, no landscape, no horizon line. No number plate and no on-screen text or lettering
anywhere at any point. Exactly six segments and exactly five hard cuts. The glossy floor carries a
full mirror reflection of the car in every segment it appears. The amber source stays lower
screen-left and the cyan source stays upper screen-right across every cut. The car is fully at
rest, static and settled, holding the hero three-quarter pack shot with clean empty negative space
upper-right, for the final second of the clip.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5 · 15s
1280x720
15.1
7,845
15.03
0
0.000
0.000
0
98
2.0 · 15s
1280x720
15.1
9,863
18.82
0
0.000
0.000
0
68
A6
HAIRLINE
Acting · two-hander, micro-expression, locked camera
The opposite of every other track: all spectacle removed so the only thing left to measure is performance. Six segments, five of them tight on a single face, the camera effectively locked off throughout. A married couple packing up after a swim, and by the end one of them has said the thing neither intended to say.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 98 cr2.0 · 68 cr1470x630 · 15.1s2 renders
Image reference 3 references · what went in
Three plates bound by job id, not URL. The location plate is deliberately sharp front to back — depth of field is the video model's job, and an earlier plate that described out-of-focus highlights as objects got them animated as grey ovals hovering over the pool in every frame. Both character sheets arrive with a completely neutral face: an expression baked into a reference becomes a mask worn for all six segments.
A 1930s open-air art-deco lido on the last afternoon of the season, photographed on 35mm at
golden hour, shot low from the concrete deck looking out across the water into a low sun. A
long rectangular pool of still teal water lined with pale turquoise and cream tiles. Cream
painted concrete deck in the foreground, wet with footprints. A low curved diving platform
and a row of folded slatted deck chairs at the far end, a rolled lane rope, a faded striped
windbreak, tall dark trees beyond the far wall.
The sun sits low and directly behind the far end of the pool, raking straight across the
water toward camera, so the surface of the water carries a fine sheet of hard sparkle. Long
raking shadows across the concrete. Fine haze in the air catching the low sun.
THE ENTIRE IMAGE IS IN SHARP FOCUS FROM THE NEAREST FOREGROUND TO THE FURTHEST TREES. Deep
focus, everything crisp, every surface resolved. The water sparkle is small, sharp and
physically on the surface of the water. Every object in the picture is a real solid thing
resting on a real surface. The air above the pool and above the deck is completely empty and
clear, containing nothing but daylight and haze.
Palette: 55% teal and turquoise water, 30% warm cream concrete and gold backlight, 15% deep
green shadow. Warm highlights against cool water, deep shadow that still holds readable
detail.
Hyperrealistic photography. Surface materials with real wear patterns, oxidization, moisture
stains, chipped and crazed paint on the tiles, salt bloom on the concrete. Water with real
reflection and refraction. Kodak Vision3 250D film emulation, visible fine film grain, soft
lens vignette, cinematic colour grade with warm mid-tones and slightly cooled shadows.
Lived-in, not pristine. Photographic, not rendered. No people, no text.
STRONG anamorphic lens character: horizontal squeeze and compression, curved barrel edge
distortion, chromatic aberration toward the left and right edges, gentle horizontal
compression of the frame. 2.39:1.
Character sheet — the woman@char_HL_woman_v1nano_banana_pro · 4k · three-view↓ Download · 5504×3072 JPEG · 2.3 MBDescriptor pasted into every video prompt @char_HL_woman_v1
female, mid-thirties, a swimmer's build, very pale skin with a hard sun line and a dense spray of dark freckles across the nose and shoulders, hair bleached almost white and cropped very short with two centimetres of dark roots, wet and slicked back, pale grey-blue eyes, a strong jaw, a fine white scar breaking the left eyebrow. Rust-orange swimsuit, a faded cream towel.
Character sheet — the man@char_HL_man_v1nano_banana_pro · 4k · three-view↓ Download · 5504×3072 JPEG · 2.6 MBDescriptor pasted into every video prompt @char_HL_man_v1
male, mid-forties, lean and weathered, deep tan with a hard pale sunglasses line across the upper cheeks, thick black hair going grey in a wide streak above the left temple, damp and pushed back, very heavy dark brows over deep-set brown eyes, a long broken nose, three days of stubble. Unbuttoned cream linen shirt, sleeves pushed to the elbow.
What the numbers say
verdict · open question · caveat
2.5 held the lock and 2.0 could not. On a byte-identical prompt asking for a locked-off tripod, 2.0 moved roughly twice as much everywhere — non-cut motion mean 5.78 against 2.5's 3.10, and a 90th percentile of 11.30 against 2.5's worst single frame of 7.20. In the final locked segment 2.0 ran 8–11 while 2.5 sat at 0.5–1.0. The close-up gap is the same story: skin-pixel share 45.0% on 2.5 against 24.8% on 2.0, whose floor drops to 8.7% — it repeatedly abandons the close-up altogether.
Does either face have a mind behind it, or is it a mask moving? Does the eyeline ladder read with the sound off? Is the break in Segment 4 earned, or does the model play the explosion from frame one?
Note the inversion against the rest of the page: here 2.0 ships the fatter file (10,613 kbps against 7,322) and more edge energy, yet loses the track outright. Bitrate is not performance.
The prompt2648 words · identical to both models
SCENE CONTEXT
Photoreal contemporary British drama, fifteen seconds, continuous in time. An open-air lido on the last afternoon of the season. A married couple are packing up after a swim. This is an ACTING SCENE and nothing else: every single segment is a tight close-up on one face, the camera is effectively locked off throughout, and everything that happens in this clip happens in two faces. There is no wide shot at any point. Both of them know something neither intends to say today, and by the end one of them has said it.
ACTIVE REFERENCES
@char_HL_woman_v1 controls her face, hair, build and swimsuit ONLY. NOT her pose, NOT her framing, NOT her expression. Female, mid-thirties, a swimmer's build, very pale skin with a hard sun line and a dense spray of dark freckles across the nose and shoulders, hair bleached almost white and cropped very short with two centimetres of dark roots, wet and slicked back, pale grey-blue eyes, a strong jaw, a fine white scar breaking the left eyebrow. Rust-orange swimsuit, a faded cream towel.
@char_HL_man_v1 controls his face, hair, build and clothing ONLY. NOT his pose, NOT his framing, NOT his expression. Male, mid-forties, lean and weathered, deep tan with a hard pale sunglasses line across the upper cheeks, thick black hair going grey in a wide streak above the left temple, damp and pushed back, very heavy dark brows over deep-set brown eyes, a long broken nose, three days of stubble. Unbuttoned cream linen shirt, sleeves pushed to the elbow.
@loc_HL_lido_v1 controls the architecture, materials, palette, lens character and direction of the light ONLY. It is an atmosphere and optics reference, not a frame to copy. NOT the framing, NOT the composition, NOT the camera position, NOT the placement of any object, and NOT any literal element of the picture. A 1930s open-air art-deco lido, off season, teal water, cream concrete, low backlit sun.
STAGING
She sits on the pool edge with her feet in the water. He is a few metres behind her on the deck, working over an open canvas bag. Both of them are present on that deck for the whole fifteen seconds and neither of them goes anywhere. The camera never shows both at once — each segment is one face — but the one who is off screen is always still there, a few metres away, and the sound tells you so. Her eyeline to him runs toward camera-right; his to her toward camera-left. The camera stays on one side of them for the whole clip so she is always framed looking camera-right and he is always framed looking camera-left. Neither of them looks into the lens.
CAMERA DISCIPLINE — READ THIS BEFORE THE SEGMENTS
The camera in this clip behaves like a locked-off tripod on a long lens. In every segment it either holds completely still or drifts a few centimetres over the entire length of the segment, slowly enough that a viewer never notices the movement while it is happening. There is one single exception, named in Segment 5, and even that is small. The camera holds one framing per segment from the first frame of that segment to the last. Every emotional escalation in this clip is carried by the performance and by the cut, never by the camera. When the scene gets bigger, the camera gets stiller.
EVENT AND SHARED DIRECTION
THE EVENT, containing both of them and never spoken aloud: the surrender of the performance.
THE SHARED DIRECTION, played by both: finish the packing up as an ordinary afternoon, and do not let the thing between them be named out loud.
THE ESSENTIAL AXIS: which of the two is willing to look at the other.
ACTING TASK — THE WOMAN (fully invested in her tactic; the work happens in her eyes)
MOTIVE: while the ordinary holds, the marriage holds. A thing not said is a thing not yet true.
GOAL: to finish packing up with him still standing there.
OBSTACLE: his flatness. Each short answer tells her he is already gone.
TACTIC: she opens warm, almost laughing, because warmth has always worked. She works him without turning round, tracking him by sound alone — the canvas, the zip, his feet on the wet concrete — and after each thing she offers she holds still and waits a beat for what comes back, then steals a look at him from the very edge of her vision. When the warmth fails she makes herself smaller and tries again. When that fails too, the thing she has been holding down all afternoon comes up through her throat before she can stop it.
Moment to moment:
— {You were good with them.} — offered light and warm, almost a laugh, her eyes down on the water while her whole attention is behind her, waiting to hear whether he takes it. The smile is real when it starts.
— the smile arrives, holds, and then dies out of her eyes first while her mouth is still doing it.
— {They believed every word.} — smaller now, the last thing she has, and on the last word her eyes flick back toward him without her head moving at all.
— {Don't. Not here.} — the break. Hurt arrives first and turns to anger inside the same second. She turns her head and not her body, and for the first time in the clip she is looking straight at him.
ACTING TASK — THE MAN (fully invested in his tactic; the work happens in his eyes)
MOTIVE: he decided before they got in the car and means to say it tomorrow, sober and indoors. This afternoon is a stay of execution he grants them both, and he knows it is cowardice under a better name.
GOAL: to get off this deck without saying it today.
OBSTACLE: her warmth. Every kind thing she offers is a bill for a performance he is about to admit was one, and he cannot take another one.
TACTIC: he keeps his hands full so his eyes have somewhere to be, and answers to objects rather than to her. He measures her from the very edge of his vision, gauging how much longer she can keep this up. Twice his eyes go to the gate at the end of the deck and come straight back.
Moment to moment:
— {Was I.} — puts her warmth back down in front of her without touching it, eyes staying on the bag. Something moves in his throat and he swallows it.
— {So did you. For nine years.} — his fuel runs out a day early. He stops, looks up at the back of her head for the first time, and hits deliberately to end it. He knows exactly what he is doing and he does it anyway.
— the half second after he says it, when it lands on him too.
— {Look at me. Once.} — everything drops away. Not an accusation, the only request he has made all day, both eyes hunting hers for any sign that she will. Then he waits, and goes on waiting, and the answer does not come, and the clip ends on him still waiting for it.
MICRO-EXPRESSION — THE WHOLE POINT OF THIS CLIP
Both faces are alive at all times and everything happens in miniature. Natural blink cadence throughout, at rest and while speaking, and a blink held a fraction too long when something lands. Eyes wet and alive with a real catchlight thrown up off the water, and the eyes always focused on something specific, never vacant, visibly moving between two points rather than staring at one. Visible breath in both chests, and a breath caught and released at each turn. The throat works: a swallow, a jaw setting, a pulse visible at the neck. Nostrils widen slightly on the intake. Micro-tension gathering and releasing at the outer corners of the eyes and between the brows. Expressions are asymmetric — one brow, one side of the mouth — and never mirrored across the face. Every reaction arrives a quarter second late, never instantly. A smile that fades leaves the eyes before it leaves the mouth. Skin flushes at the throat and the tops of the ears under pressure. Small involuntary movements: a finger tightening on cloth, a shoulder dropping, a chin lifting a few millimetres. The eyes fill without anything falling.
FORMAT MODE
Six segments with HARD CUTS. Every segment is a tight close-up on one face. Real time, no ramps, no speed changes. Dialogue in British English. Sound design and location tone only, no music, no voice-over.
SEGMENT 1 — HER, CLOSE (0.0s–2.6s)
Tight on THE WOMAN, head and shoulders filling the frame, three-quarters away from camera, the low sun full on the side of her face. Her hands are working the towel at the bottom of frame. She speaks out across the water, warm, almost laughing.
LENS: locked off and completely still. Widescreen, very shallow, backlit, clean crisp glass with a matte finish.
THE WOMAN {You were good with them.}
SFX: <water moving against tile, close> <a wet towel wrung out> <canvas shifting a few metres behind her>
HARD CUT.SEGMENT 2 — HIM, CLOSE (2.6s–5.2s)
Tight on THE MAN, head and shoulders filling the frame, backlit down one edge, his eyes down on the bag. He answers to the bag. Something moves in his throat afterwards.
LENS: locked off and completely still. Widescreen, very shallow, clean crisp glass with a matte finish.
THE MAN {Was I.}
SFX: <canvas shifting> <a zip half pulled and left>
HARD CUT.SEGMENT 3 — HER, BIG CLOSE (5.2s–7.6s)
Very tight on THE WOMAN, her eyes filling the upper half of frame, the flat teal water going soft far below her. Her smile is still on her mouth and already gone from her eyes. She offers the last thing she has to the empty water in front of her, and on the last word her eyes flick back toward him without her head moving at all.
LENS: locked off. At most the frame drifts a couple of centimetres closer across the whole segment, too slowly to be noticed. Widescreen, extremely shallow, clean crisp glass.
THE WOMAN {They believed every word.}
SFX: <the water going flat> <a breath through the nose>
HARD CUT.SEGMENT 4 — HIM, CLOSE (7.6s–10.2s)
Tight on THE MAN. He stops working. He raises his head and looks at the back of her head, the first time in the clip he has looked at her at all, and hits deliberately. Then it lands on him too and he holds still with it.
LENS: locked off and completely still. Widescreen, very shallow, clean crisp glass.
THE MAN {So did you. For nine years.}
SFX: <a hand going still on canvas> <wind across an empty pool>
HARD CUT.SEGMENT 5 — HER, BIG CLOSE (10.2s–12.6s)
Very tight on THE WOMAN, the low sun full on her face. Her hands stop dead. She turns her head over her right shoulder toward him, her head only and not her body, and a whole afternoon of holding arrives at once. Hurt first, then anger, inside the same second. This is the largest thing that happens in the clip and it happens entirely in her face.
LENS: the one exception — the faintest handheld tremor enters the frame, the size of a held breath, and the framing itself does not change. No push, no move. Widescreen, extremely shallow, clean crisp glass.
THE WOMAN {Don't. Not here.}
SFX: <a breath taken hard> <all location tone dropping away for half a second>
HARD CUT.SEGMENT 6 — HIM, CLOSE, AND HELD (12.6s–15.0s)
Tight on THE MAN, the same framing he had in Segment 4, unchanged. Everything drops out of his face and he asks. Then he waits. He goes on waiting. Nothing else happens: no move, no turn, no answer, no cut. His face does the last two seconds of this film on its own, and the clip ends with him still waiting and the answer still not coming.
LENS: locked off, completely motionless, identical framing from the first frame of the segment to the last. Widescreen, very shallow, backlit, clean crisp glass with a matte finish.
THE MAN {Look at me. Once.}
SFX: <a deck chair creaking once, a few metres away> <wind> <held silence to the last frame>
DIALOGUE
British English. Both voice descriptors used exactly as written.
THE WOMAN — voice: warm, low English female, careful and evenly paced, consonants softened by tiredness. It starts light and almost laughing, it gets smaller, and on the break the pitch rises once and cracks at the top of it.
THE MAN — voice: dry mid-range English male, educated and unhurried, ends of lines dropping almost to breath. Flat throughout, and on the last line completely emptied out and quiet.
Twenty-four words across fifteen seconds. The silences are the scene, and there is dead air before the last line and dead air after it.
PHYSICS
Still water with real surface tension: her feet move in it and the rings spread and die away, and the broken light on the surface scatters where it is disturbed and settles again as the water stills. The cream towel is soaked and heavy. Wet skin dries unevenly in the sun. The surface of the pool carries the low sun and the open sky and nothing else, and the light bouncing up off it throws a soft moving net of brightness under both faces.
LIGHTING
One source and one bounce. The sun is low and directly behind the far end of the pool, raking across the water into the lens, so both faces are edge-lit. The cream concrete deck bounces a soft warm fill into the shadow side of both faces so the eyes always hold detail. Every light source renders as a soft, contained glow within its own source.
STYLE
Fully photoreal contemporary drama, 35mm large-format digital cinema capture on a long lens, widescreen 2.39:1, clean crisp lens glass with a matte finish, extremely shallow depth of field throughout, fine organic film grain, warm golden backlight against deep teal water, a restrained naturalistic grade. Still, patient, unshowy camerawork of the kind used when the performance is the only thing being sold.
POSITIVE LOCKS
Every one of the six segments is a tight close-up of a single face filling most of the frame. The clip contains no wide shot and no two-shot at any point.
The camera holds one fixed framing for the whole of each segment and every move in the clip is small enough to go unnoticed while it happens. The scene escalates through the faces and the cuts alone.
Exactly two people exist in this scene, matching @char_HL_woman_v1 and @char_HL_man_v1 exactly, same faces, hair and wardrobe throughout, and both of them remain present on the deck for the entire fifteen seconds.
Every human figure in every frame is one whole real body, and the surface of the water carries only sky, the low sun and moving highlights.
Every out-of-focus area of every frame is a smooth continuous gradient of light and colour, and every object in the frame is a real physical thing resting on a real surface.
The air above the pool and the deck is completely empty and clear in every frame.
Both faces carry continuous living micro-expression: focused wet eyes with a real catchlight, natural blinking, visible breath, a working throat and small asymmetric movements of brow, jaw and mouth in every segment in which they appear.
She is always framed looking toward camera-right and he is always framed looking toward camera-left, and the camera stays on one side of them throughout.
Exactly six segments and exactly five hard cuts.
The final segment is a motionless close-up of THE MAN that holds unchanged to the very last frame, with him waiting and no answer arriving.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1470x630
15.1
7,322
14.05
19
1.050
1.062
42.49
98
2.0
1470x630
15.1
10,613
20.23
32
1.042
1.056
44.58
68
A6b
HAIRLINE ✚
Scene extension · 2.5 only, 2.0 has no equivalent
The one capability on this page that is not a quality comparison. 2.5 can continue an existing clip forward from its last frame; 2.0 exposes no extension mode at all, so there is nothing to put beside it. The wipe here runs source against continuation rather than model against model — drag it to look for the join.
The bound reference is a video, not an image — the job id of the clip being continued, with mode: "video_extension" and extension_mode: "forward". Segment 1 reproduces the source's final framing exactly and holds it before anything moves, and POSITIVE LOCKS asserts that the first frame of this clip is the last frame of the reference video.
What the numbers say
verdict · open question · caveat
The join is optically invisible. Measured against the source's final frame it lands at the same level as an ordinary in-shot frame pair — far below the >38 a hard cut scores in the same measurement. That only worked because the source was written to end motionless: an earlier attempt off a clip ending on a moving pull-back produced a visible discontinuity and dropped a character. Same model, same settings — the only variable was how the source ended. A second arm tested whether image references help: identical prompt and source, one run bound to the video only, one with the two character sheets and the location plate added alongside. Every metric landed within noise. video_extension reads identity, wardrobe, grade and lens off the source clip's own frames, so the images have nothing left to contribute — if an extension looks wrong, the fault is in the source's final frames or the prompt, never a missing image reference.
Does it feel like the same scene continuing, or like a sequel shot on a different day? Is the seam invisible enough to cut through, or does it need a hard cut to hide it? Does the four-word version out-act the twenty-eight-word version?
Aspect ratio follows the source video, not the request — 21:9 was asked for in both directions and the output matched the source each time. If you want a 2.39:1 extension, the source must already be 2.39:1. The seam figure is quoted as a relationship rather than an absolute on purpose: it moves with the downscale used to compute it, so only the comparison against a same-shot pair and against a real cut is stable.
The prompt2505 words · identical to both models
CONTINUATION CONTEXT
This is the fifteen seconds immediately AFTER the reference video, running forward from its last
frame with no gap in time. Same afternoon, same empty off-season lido, same two people, same
wardrobe, same low backlit sun, same lens character, same grade. The reference video ends on a
motionless close-up of THE MAN, who has just asked {Look at me. Once.} and is still waiting for an
answer that has not come. This clip is the answer coming. Like the reference video it is an ACTING
CLIP and nothing else: every segment is a tight close-up on one face, there is no wide shot at any
point, and the camera is effectively locked off throughout.
Photoreal contemporary British drama, restrained, real time, continuous.
ACTIVE REFERENCES
@char_HL_woman_v1 controls her face, hair, build and swimsuit ONLY. NOT her pose, NOT her framing,
NOT her expression. Female, mid-thirties, a swimmer's build, very pale skin with a hard sun line and
a dense spray of dark freckles across the nose and shoulders, hair bleached almost white and cropped
very short with two centimetres of dark roots, wet and slicked back, pale grey-blue eyes, a strong
jaw, a fine white scar breaking the left eyebrow. Rust-orange swimsuit, a faded cream towel. 100%
matches the reference video.
@char_HL_man_v1 controls his face, hair, build and clothing ONLY. NOT his pose, NOT his framing, NOT
his expression. Male, mid-forties, lean and weathered, deep tan with a hard pale sunglasses line
across the upper cheeks, thick black hair going grey in a wide streak above the left temple, damp and
pushed back, very heavy dark brows over deep-set brown eyes, a long broken nose, three days of
stubble. Unbuttoned cream linen shirt, sleeves pushed to the elbow. 100% matches the reference video.
CAMERA DISCIPLINE — READ THIS BEFORE THE SEGMENTS
The camera behaves like a locked-off tripod on a long lens, exactly as in the reference video. In
every segment it either holds completely still or drifts a couple of centimetres across the entire
segment, slowly enough that a viewer never notices it happening. It holds one framing per segment
from that segment's first frame to its last. Everything in this clip is carried by two faces and by
the cuts. When the scene gets bigger, the camera gets stiller.
STAGING
Exactly as the reference video leaves them. She is on the pool edge with her feet in the water,
screen-left. He is behind her on the deck by the open canvas bag, screen-right. The camera stays on
the same side of them it has been on the whole time, so she remains the screen-left person and he
remains the screen-right person. Neither of them looks into the lens. The distance between them does
not close.
EVENT AND SHARED DIRECTION
THE EVENT, containing both of them and never spoken aloud: the moment after the thing is true, when
there is nothing left to protect and both of them go on packing anyway.
THE SHARED DIRECTION, played by both: get through the next fifteen seconds without either of them
having to say the word for what has just happened.
THE ESSENTIAL AXIS: which of the two is willing to look at the other. In the reference video she
would not. In this clip she does, and it is worse.
ACTING TASK — THE WOMAN (fully invested in her tactic; the work happens in her eyes)
MOTIVE: the performance is over and she is the one who ended it, and she needs to know exactly what
she has ended before she gets up off this ledge.
GOAL: to look at him properly, once, and read the whole thing off his face.
OBSTACLE: her own body — it will not hold still and her breath keeps going out of time, and she can
feel the crying getting organised behind her eyes and refuses it the room.
TACTIC: she turns and takes him in deliberately, the way you look at something you are memorising.
She goes over his face inch by inch — both eyes first, then the mouth, then back to the eyes — and
what she is hunting for is a single sign he does not want this. She does not find it. She lets go of
the towel. Then she puts herself back together in front of him, on purpose, so that he watches her
do it.
Moment to moment:
— the look, given at last and held far past comfortable. The eyes going over his face in small
purposeful moves, checking, checking again.
— {There it is.} — quiet, almost kind, delivered while she is still reading him.
— she stops looking. Something goes out of her shoulders and does not come back.
— she wipes her face once with the back of her wrist, briskly, the way you deal with water, and goes
back to her feet in the pool.
ACTING TASK — THE MAN (fully invested in his tactic; the work happens in his eyes)
MOTIVE: he asked for the look and he got it, and he has spent nine years being the one who is looked
at like this and never once until now been unable to answer it.
GOAL: to stay in front of her for as long as she needs to look, and not look away first.
OBSTACLE: she is right, and he can feel his own face agreeing with her while he tries to hold it
still.
TACTIC: he takes the look. He does not defend, does not explain, does not fill the silence. He
watches her working over his face and lets her, and when he sees her reach the end of it he stops
holding his own face together. When she starts to move he goes back to the bag, because giving her
somewhere to not be looked at is the last kind thing available to him.
Moment to moment:
— receiving the look, both his eyes going between both of hers, staying.
— {I know.} — said to her and not to the bag, the first time all afternoon.
— the half second after she says {There it is.}, when his face stops working.
— he breaks first — not away, but down — and picks up the bag because his hands need it.
MICRO-EXPRESSION — THE WHOLE POINT OF THIS CLIP
Both faces are alive at all times and everything happens in miniature, smaller here than in the
reference video because there is nothing left to perform. Natural blink cadence throughout, at rest
and while speaking, with a blink held a fraction too long each time something lands, and a run of
fast blinks when either of them is refusing to cry. Eyes wet and alive with a real catchlight thrown
up off the water, always focused on something specific, never vacant, and visibly moving between the
partner's two eyes rather than staring at one point. Visible breath in both chests, uneven, caught
and released at each turn, one breath going in and not coming back out for a beat. The throat works:
a swallow, a jaw setting and releasing, a pulse visible at the side of the neck. Nostrils widen
slightly on the intake. Micro-tension gathering and releasing at the outer corners of the eyes,
between the brows, and in the small muscle at the corner of the mouth. Expressions asymmetric — one
brow, one side of the mouth — and never mirrored across the face. Every reaction arrives a quarter
second late, never instantly. The chin dimples and stills. Skin flushes at the throat, the tops of
the ears and the outer edge of the nose. The eyes fill without anything falling. Small involuntary
movements: a finger tightening on wet cloth and letting go, a shoulder dropping, a swallow that does
not go down cleanly.
FORMAT MODE
Six segments with HARD CUTS. The first is the only wide; the other five are tight on a single face.
Real time, no ramps. British English. Sound design and location tone only, no music, no voice-over.
SEGMENT 1 — HIM, CLOSE, PICKING UP EXACTLY WHERE IT LEFT OFF (0.0s–2.4s)
The identical close-up of THE MAN the reference video ends on, same framing, same light, held
unchanged. He is still waiting for the answer. He does not speak. The wait goes on a beat past
bearable, and then something crosses his face that says he has stopped expecting it — and in the
same moment he hears her move behind him, and his eyes come up.
LENS: locked off, completely motionless, identical framing to the last frame of the reference video.
Widescreen, very shallow, backlit, clean crisp glass with a matte finish.
SFX: <wind across an empty pool> <water displaced behind him as she turns>
HARD CUT.SEGMENT 2 — HER, BIG CLOSE, THE LOOK (2.4s–5.4s)
Very tight on THE WOMAN, her face filling the frame, the low sun full on it, her eyes wet and hard
lit. She is looking straight at him and going over his face in small purposeful moves. She says
nothing. She is reading.
LENS: locked off. The smallest handheld float, the size of a held breath, and the framing does not
change. Widescreen, extremely shallow, clean crisp glass.
SFX: <a breath going in and not coming back out> <distant wind>
HARD CUT.SEGMENT 3 — HIM, BIG CLOSE, TAKING IT (5.4s–8.0s)
Very tight on THE MAN, backlit down one edge, warm bounce off the concrete filling the shadow side so
his eyes hold detail. He takes the look. His eyes go between both of hers and stay. He answers her
without being asked anything.
LENS: locked. Widescreen, extremely shallow, clean crisp glass.
THE MAN {I know.}
SFX: <a swallow> <the canvas bag settling>
HARD CUT.SEGMENT 4 — HER, BIG CLOSE, THE VERDICT (8.0s–10.6s)
Very tight on THE WOMAN, still on him, reaching the end of what she is reading. She says it almost
kindly, still looking. Then she stops looking, and something goes out of her shoulders.
LENS: locked off. At most a couple of centimetres of drift across the whole segment, too slow to be
noticed. Widescreen, extremely shallow, clean crisp glass.
THE WOMAN {There it is.}
SFX: <water displaced by a foot> <all location tone thinning for half a second>
HARD CUT.SEGMENT 5 — HIM, CLOSE, THE FACE STOPPING (10.6s–12.8s)
Tight on THE MAN in the half second after it lands and the seconds after that. Nothing is said. His
face stops holding itself together, not into a big expression but into no expression at all, which is
worse. Then he reaches down out of frame for the bag because his hands need something.
LENS: locked off and completely still. He reaches down out of the bottom of the frame; the camera
does not follow him. Widescreen, very shallow, clean crisp glass.
SFX: <a zip pulled all the way> <a bag lifted off wet concrete>
HARD CUT.SEGMENT 6 — HER, CLOSE, PUTTING IT BACK (12.8s–15.0s)
Tight on THE WOMAN, turned back to the water, in profile to three-quarters. She wipes her face once
with the back of her wrist, briskly, the way you deal with water. Her feet move in the pool and the
rings go out. She is composed by the last frame and it is not the same composure she had at the
start. She holds still and the clip ends on her.
LENS: locked off completely, no movement at all, coming to rest for the final second. Widescreen,
shallow, backlit, clean crisp glass with a matte finish.
SFX: <a wrist against a wet cheek> <rings of water spreading> <wind> <held to the last frame>
DIALOGUE
British English. Both voice descriptors used exactly as written.
THE WOMAN — voice: warm, low English female, careful and evenly paced, consonants softened by
tiredness. Here it is quiet and steady and slightly wrecked underneath, and it does not rise.
THE MAN — voice: dry mid-range English male, educated and unhurried, ends of lines dropping almost to
breath. Completely emptied out, barely above a breath.
Five words across fifteen seconds. The silence is the scene. No voice-over of any kind.
PHYSICS
Still water with real surface tension: her feet move in it, the rings spread and die away, and the
reflection distorts where the surface is disturbed and reforms as it stills. The cream towel is
soaked and heavy and holds its folds. Wet skin dries unevenly in the sun. The canvas bag has real
weight when it is lifted. The water returns a true optical reflection of the deck, correct positions
and correct reflected handedness.
LIGHTING
Identical to the reference video. One source and one bounce. The sun is low and directly behind the
far end of the pool, raking across the water into the lens, so both faces are edge-lit. The cream
concrete deck bounces a soft warm fill into the shadow side of both faces so the eyes always hold
detail. Every light source renders as a soft, contained glow within its own source. The sun has moved
a few degrees lower across the fifteen seconds and nothing else has changed.
STYLE
Fully photoreal contemporary drama, 35mm large-format digital cinema capture, widescreen 2.39:1,
clean crisp lens glass with a matte finish, extremely shallow depth of field on the five close
segments, fine organic film grain, warm golden backlight against deep teal water, a restrained
naturalistic grade. Identical grade, grain, lens character and colour to the reference video.
POSITIVE LOCKS
The first frame of this clip is the last frame of the reference video — the same motionless close-up
of THE MAN, same framing, same light, same grade, same grain — and the join is invisible.
Every one of the six segments is a tight close-up of a single face filling most of the frame. The
clip contains no wide shot and no two-shot at any point.
The camera holds one fixed framing for the whole of each segment and every move in the clip is small
enough to go unnoticed while it happens.
Both people remain present on that deck for the entire fifteen seconds and neither of them goes
anywhere or leaves the scene.
Exactly two people are present, matching @char_HL_woman_v1 and @char_HL_man_v1 and the reference
video exactly.
Every out-of-focus area of every frame is a smooth continuous gradient of light and colour, and every
object in the frame is a real physical thing resting on a real surface.
The air above the pool and the deck is completely empty and clear in every frame.
Both faces carry continuous living micro-expression: focused wet eyes with a real catchlight, natural
blinking, visible breath, a working throat and small asymmetric movements of brow, jaw and mouth in
every segment in which they appear.
She is the screen-left person and he is the screen-right person, and the camera stays on the same
side of them it was on in the reference video.
Exactly six segments and exactly five hard cuts.
The clip ends locked off on her face, composed and still, with the water settling.
The numbers3 renders · bitrate, detail, macroblocking
extension + image refs
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
source · 2.5
1470x630
15.1
7,322
14.05
19
1.050
1.062
42.49
98
extension · 2.5
1470x630
15
6,802
13.03
12
1.056
1.068
46.55
98
extension + image refs
1470x630
15
6,456
12.38
12
1.062
1.074
46.91
98
A7
REDLINE
Music video · maximum chaos, storyboard-bound, match transitions
The exact opposite of HAIRLINE. That track stripped out all spectacle to test acting; this one strips out all story to test spectacle — twelve beats, eleven liquid morphs, violent camera on every one, zero dialogue. It is also the only track on this page authored storyboard-first: a 12-panel board was generated and measured before a single frame of video was bought.
⇆
2.52.0
Drag the handle · same prompt, byte for byte
2.5 · 98 cr2.0 · 68 cr1470x630 · 15.1s2 renders
Image reference — storyboard 2 references · what went in
A grid-shaped reference gets copied as a grid unless you exclude the layout twice — once in a connector block beside the reference, once again in the positive locks. The board controls palette, smear language, light arc and beat order only; framing, composition and panel layout are explicitly excluded both times.
A twelve-panel film storyboard laid out as a clean 4x3 grid, four panels across and three
panels down, thin gutters, every panel the same widescreen 2.39:1 letterbox shape. No
numbers, no text, no lettering, no arrows, no annotations anywhere.
RULE ONE: EVERY SINGLE PANEL IS CAUGHT MID-MOVEMENT, with heavy directional motion blur
running through it. Nothing is posed, nothing is still. If a panel could be a photograph of
something standing still, it is wrong.
RULE TWO: one unbroken molten orange-red ribbon of light runs through ALL TWELVE PANELS,
clearly visible in every one, entering one edge and leaving the other, reading as a single
continuous line of light travelling across the whole grid.
RULE THREE: this is a film about people. Human faces, human hands and human bodies appear in
at least half the panels, always in motion, always smeared, never sharp and never posed.
STYLE: extreme long-exposure motion smear, deep black silhouettes, cold teal as the only
cool colour, heavy real film grain, strong bloom, chromatic fringing. Light arc: top row
near-black night, middle row detonating into white, bottom row blown-out cream. Reading
order left to right, top row then middle then bottom.
TOP ROW, night: 1 extreme macro of a human eye, orange ribbon across the wet surface, frame
rushing inward. 2 the pupil as a black circle with teal corona, streaking inward. 3 the rim
of an enormous pale moon breaking into frame, ribbon bending around it. 4 a low wide sports
car in black silhouette streaking across the moon, smeared horizontal, the ribbon its
tail-light trail.
MIDDLE ROW, the crossover: 5 extreme macro of a spinning wheel dissolved into a solid disc
of blur, orange glowing within. 6 a single human face in close-up caught mid-shout, mouth
wide open, head thrown back, smeared sideways, teal rim light along the jaw. 7 a mass of
raised hands and forearms in black silhouette, all blurred, teal trails between them. 8 two
human bodies caught mid-jump in the air, limbs streaked, frame blown almost fully white.
BOTTOM ROW, blown-out cream: 9 a hard vertical edge of white light wiping across the frame
from left to right, half the image already white. 10 long vertical folds of streaming fabric
in hard wind, feathered into teal bands across cream, a human silhouette bowed inside them.
11 a single human hand reaching through a gap in the fabric toward camera, fingers blurred
with speed. 12 a curved expanse of glossy black car bodywork sweeping past camera, and held
in its reflection the same human eye from panel one, the ribbon terminating in its pupil.
Hyperrealistic photography, real long-exposure motion blur, real film grain, real bloom.
Photographic, not illustrated, not a drawing. No text, no watermark, no logos, no panel
numbers, no borders around the outside of the grid.
What the numbers say
verdict · open question · caveat
The clearest split on the page, and it inverts the usual direction. Asked for eleven liquid morphs and no cuts at all, 2.5 produced 6 detected discontinuities against 2.0's 15 — 2.5 obeyed the morph instruction where 2.0 fell back on cutting. An earlier version measured the same thing from the other side: 2.5 hit 11 of 11 specified transitions, 2.0 hit 4. And 2.5 is the only render anywhere in this episode to measure below the macroblocking floor, at 0.963 against 2.0's 1.065 — on the most compression-hostile material in the whole test.
Does it read as one continuous move, or as separate clips bolted together? Which model's speed ramp actually lands? Does the eye → moon → paint loop close, or does the model forget the eye before the end?
The board is the reason it moves. An earlier six-panel board gave the model ~2.5s per beat and it idled between them — the fix was upstream of the video prompt entirely. At twelve panels each beat gets ~1.25s and it never settles. Row luma on the shipped board climbs 50.4 → 99.6 → 146.0, monotonic; the 15-panel alternative rises by row too but its per-panel arc oscillates badly, because at five across in a 4k render GPT starts losing per-panel detail.
The prompt1597 words · identical to both models
SCENE CONTEXT
Fifteen seconds of pure adrenaline. Twelve beats, roughly one and a quarter seconds each. Cars, bodies, faces, cloth and light, all at maximum speed, all melting into each other. No story. No dialogue. Only momentum and light.
RULE ONE — NOTHING EVER STOPS
Not one still frame. Not one held pose. Not one moment of rest, from the first frame to the last. The camera is moving continuously for the entire fifteen seconds and never parks, never settles, never holds. Every subject is caught mid-movement with heavy directional motion blur. Every element in the frame is travelling: the foreground travels, the background travels, the light travels, the grain travels. If any frame could pass for a photograph of something standing still, it is wrong.
RULE TWO — EVERYTHING IS LIQUID
There are NO CUTS in this clip. Not one. Every one of the eleven transitions is a LIQUID MORPH: the outgoing image stretches, smears, melts and pours itself into the incoming one while the camera keeps travelling straight through. Shapes flow into other shapes. Nothing snaps, nothing steps, nothing jumps, nothing blinks to black. Treat the whole fifteen seconds as one continuous elastic material being pulled through twelve different shapes.
RULE THREE — THE RIBBON IS ALIVE
There is one molten orange-red light form travelling through the clip and it BREATHES. It is never a straight line. It is never an even bar. It is never the same width for more than a few frames. It swells and collapses: it blooms out fat and blinding across a third of the frame, then starves down to a hairline thread, then vanishes completely for a beat and slams back bigger and hotter. It whips, coils, snakes, doubles back on itself, tears into separate strands and rejoins them. It bends around everything it passes. Its brightness pumps continuously between almost invisible and pure white-hot. IF IT EVER READS AS A STEADY EVEN STRIPE LYING FLAT ACROSS THE FRAME, IT IS WRONG.
ACTIVE REFERENCE
@style_RL_board_v1: the attached image is a STYLE AND BEAT reference ONLY. Take from it: the palette, the extreme long-exposure motion smear, the black silhouettes, the teal and orange colour language, the heavy grain and bloom, the dark-to-bright light arc, and the order of the twelve beats. Take from it NOTHING of its layout. The output is one single continuous full-frame film with no grid, no panels, no gutters, no dividing lines, and no part of the frame ever split into sections.
FORMAT MODE
Twelve beats, eleven liquid morph transitions, fifteen seconds, one unbroken accelerating movement. Sound design only. No text, no captions, no logos. No readable identities.
BEAT 1 (0.0–1.25s) NIGHT. Extreme macro of a human eye wide open, iris burning teal. The ribbon is a thin hot filament crawling across the wet surface, pulsing brighter. CAMERA: crash push straight at the pupil, accelerating hard.
The pupil floods outward and swallows the frame—
BEAT 2 (1.25–2.5s) Falling through blackness, teal streaks tearing past on both sides. The ribbon spirals ahead of the camera, thickening and thinning. CAMERA: hurtling forward with a slow roll.
The spiral opens into a circle of light—
BEAT 3 (2.5–3.75s) An enormous pale moon low in a black sky. The ribbon whips around its rim and snaps away. CAMERA: fast rise up the face of it.
The rim stretches sideways into a road—
BEAT 4 (3.75–5.0s) A low wide sports car in pure black silhouette tearing across the moon's face, smeared into a long horizontal blur. The ribbon explodes off its tail lights, fat and blinding. CAMERA: violent whip pan tracking with the car, background shredding into horizontal streaks.
The tail-light bloom collapses into a spinning disc—
BEAT 5 (5.0–6.25s) Extreme macro of a spinning wheel, spokes dissolved into a solid disc of blur, orange burning through from inside it. The ribbon coils around the rim. CAMERA: rotating with the wheel and pulling out fast.
The disc tears open and floods the frame with white—
BEAT 6 (6.25–7.5s) A human face in tight close-up caught mid-scream, mouth wide open, head thrown back, teal rim light along the jaw, the whole face smeared sideways by the speed of the camera. The scream is raw and full-throated. The ribbon vanishes completely here. CAMERA: ripping past the face left to right.
The open mouth stretches and becomes a tunnel—
BEAT 7 (7.5–8.75s) A second face beside it, eyes screwed shut, chin up, dissolving into white as the frame blows out. The ribbon slams back twice as thick. CAMERA: continuing the same rip, faster.
The white bleeds apart into bodies—
BEAT 8 (8.75–10.0s) A mass of raised hands and forearms in black silhouette, all blurred, teal light trails threading between them. The ribbon snakes through the gaps between arms. CAMERA: driving straight through the middle of them.
The arms elongate and lift off the ground—
BEAT 9 (10.0–11.25s) Two human bodies caught mid-jump in the air, limbs streaked into ribbons, the frame almost fully white. The ribbon strobes on and off. CAMERA: rising underneath them.
A blade of white light pours across the frame and washes them away—
BEAT 10 (11.25–12.5s) Long vertical folds of streaming fabric in hard wind, feathered into deep teal bands across blown-out cream, a human silhouette bowed inside them. The ribbon burns through the torn gaps. CAMERA: craning up through the cloth.
The cloth peels back off something moving underneath—
BEAT 11 (12.5–13.75s) A single human hand bursting through a gap in the fabric straight toward camera, fingers streaked with speed. The ribbon wraps the wrist and rips away. CAMERA: pulling back fast ahead of the hand.
The hand's motion smear pours into a curve of black—
BEAT 12 (13.75–15.0s) A curved expanse of glossy black car bodywork sweeping past camera, teal raking across the paint, and held in its reflection the same human eye from the first beat. The ribbon narrows to a single hot filament and drives into the pupil. CAMERA: still travelling, still sweeping, right to the final frame. The clip ends mid-move.
PHYSICS
Everything smears along its own direction of travel, never in a random direction. The car's blur runs horizontal because the car runs horizontal. The cloth's blur runs vertical because the cloth falls vertical. Limbs smear along the arc they swing through. Light trails are continuous and unbroken, never dotted or stuttering. Bloom spreads outward from bright areas into the dark, never the reverse.
LIGHTING
One hot orange-red source that is the ribbon itself, and one cold teal source rimming everything else. Nothing is lit by anything but those two. Beats 1 to 5 sit in near-total black. Beats 6 to 9 are a rising white blowout. Beats 10 to 12 are blown-out cream with black silhouettes inside them. Every light source blooms outward from its own core.
AUDIO — PURE SOUND DESIGN, ABSOLUTELY NO MUSIC
The entire soundtrack is designed sound effects and room tone. There is NO music of any kind: no beat, no drums, no percussion loop, no bass line, no melody, no chords, no instruments, no synth pads, no score, no rhythm track. There is NO voice of any kind except one wordless scream: no voice-over, no narration, no dialogue, no singing, no lyrics, no chanting, no crowd chanting, no whispering.
The sound is built only from these: <a deep sub-bass pressure swell rising from silence and building through the first three beats> <a low metallic resonance as the pupil opens> <air tearing past the lens through the fall> <a hard engine doppler screaming past from left to right> <tyre roar and the mechanical whine of a wheel at speed> <one raw wordless human scream with a long reverb tail, and nothing under it> <all sound cutting to absolute silence for a quarter second at the white blowout> <a single enormous low impact returning out of that silence> <a filtered noise sweep rising as the bodies lift> <hard wind and heavy fabric snapping like a whip> <the deep sub-bass pressure still running, unresolved, at the final frame>.
STYLE
Hyper-real abstract advertising photography, extreme long-exposure motion smear, subjects dissolving into streaks of their own movement, deep black silhouettes with no interior detail, widescreen 2.39:1. Teal and orange are the only two colours. Heavy real film grain, violent bloom, chromatic fringing toward the left and right edges. Photographic and shot, never illustrated and never rendered.
DO NOT INCLUDE, UNDER ANY CIRCUMSTANCES
No music, no beat, no drums, no percussion, no bass line, no melody, no instruments, no synth, no score, no soundtrack. No voice-over, no narration, no dialogue, no speech, no singing, no lyrics, no chanting. No text, no captions, no subtitles, no titles, no lettering, no numbers, no logos, no watermarks, no number plates. No grid, no panels, no gutters, no dividing lines, no split screen, no picture-in-picture, no borders. No freeze frames, no locked-off static shots, no held poses, no settling, no stopping, no black frames, no fade to black, no end card.
THE LAST WORD — READ THIS TWICE
The single most important quality of this clip is UNBROKEN MOTION. Every element on screen is moving at all times and in every direction: things rush toward the lens, away from it, across it, up through it and down past it. Foreground elements streak past while background elements sweep the other way. The light itself is always travelling. The camera never once holds still, never once settles, and never once arrives. There is no moment anywhere in these fifteen seconds where the picture rests. It begins mid-movement, it stays in movement, and it ends mid-movement with the camera still travelling.
The numbers2 renders · bitrate, detail, macroblocking
Render
Resolution
Dur
kbps
MB
Detail
Block
Block graded
Δframe
Credits
2.5
1470x630
15.1
13,208
25.11
15
0.963
0.926
133.15
98
2.0
1470x630
15.1
15,208
28.87
22
1.065
1.097
101.81
68
04 — the raw table
Every measurement
All 22 rendersbitrate, detail, shadow, macroblocking, motion, cost
Clip
Model
Resolution
Dur
MB
kbps
Detail
Shdw
Block
Block graded
Δframe
Credits
T1_25
2.5
1280x720
8.1
3.11
2,933
107
64
1.065
1.051
13.37
52
T1_20
2.0
1280x720
8.1
4.65
4,460
144
65
1.045
1.053
26.6
36
T2_25
2.5
1280x720
8.1
4.9
4,715
169
63
1.040
1.058
13.15
52
T2_20
2.0
1280x720
8.1
4.46
4,279
151
65
1.042
1.050
24.94
36
T3_25
2.5
1280x720
8.1
2.4
2,224
193
65
1.079
1.083
5.85
52
T3_20
2.0
1280x720
8.1
3.88
3,702
90
65
1.073
1.114
26.91
36
T3_25hi
2.5 high-bitrate
1280x720
8
11.52
11,301
102
64
1.253
1.314
23.89
52
T3_20hi
2.0 high-bitrate
1280x720
8.1
10.61
10,398
165
64
1.080
1.099
29.23
36
T3_20_1080
2.0 1080p
1920x1080
8.1
7.8
7,595
147
65
1.039
1.055
35.51
72
T3_20_4k
2.0 4K
3840x2160
8
5.22
5,035
118
64
1.004
1.047
19.23
176
T4_25
2.5
1280x720
8
3.78
3,597
511
65
1.012
1.030
21.84
52
T4_20
2.0
1280x720
8.1
3.9
3,718
263
65
1.020
1.029
31.02
36
T5_25
2.5
1280x720
8.1
0.81
645
137
63
1.096
0.870
5.34
52
T5_20
2.0
1280x720
8.1
2.81
2,632
161
62
1.349
1.784
6.25
36
T6_25
2.5
1280x720
20.1
8.23
3,141
120
63
1.103
1.082
22.9
130
T6_20
2.0
1280x720
15.1
6.85
3,492
58
65
1.109
1.102
22.72
68
T5_25hi
2.5 high-bitrate
1280x720
8.1
4.79
4,601
112
65
1.015
1.188
7.2
52
T5_20hi
2.0 high-bitrate
1280x720
8.1
7.21
7,018
306
64
1.000
0.740
4.38
36
T7_25
2.5 omni_reference
1280x720
8.1
2.07
1,895
37
62
1.090
1.094
9.81
52
T7_20
2.0 image_references
1280x720
8.1
3.17
2,989
44
65
1.034
1.061
17.83
36
T5_omni
Gemini Omni Flash
1280x720
8
2.18
2,035
81
48
1.134
0.947
8.28
24
T5_cs30
Cinema Studio 3.0
1280x720
8.1
1.74
1,587
80
64
1.319
1.160
3.12
40
Detail — Laplacian variance over 5 evenly-spaced frames, all normalised to 1280×720 first. Higher means more real edge and texture information. Shdw — distinct luma levels below Y=64. A low count means posterisation. Everything here sits at 57–65 of ~64 possible, so treat small differences as noise. Block — gradient energy on 8-pixel block boundaries divided by interior gradient energy. 1.000 is the floor: no detectable macroblocking. Only comparable within one output resolution. Block graded — the same measure after contrast ×1.55, shadow lift to gamma 0.82, saturation ×1.35. Δframe — mean absolute luma delta between consecutive sampled frames. A motion-energy proxy, not a quality score.
05 — what to take care of
Learnings
The things that cost time, or would have. Each one traces back to a measurement on this
page rather than a hunch.
01always on
The default bitrate is the bug
bitrate_mode ships as "standard" and costs exactly the same as "high" — preflighted on every call, identical to the credit. On the TITLE track the flag took macroblocking from 1.349 to 1.000 (the floor, no detectable block structure) and nearly doubled detail. Set it on every Seedance call you make.
02method
There is no seed control, so look is never isolated
Higgsfield exposes no seed. Every call is a fresh generation, so a "standard vs high" pair differs in content as well as encoding. Container facts — bitrate, file size, block structure — are hard. Any claim about which one looks better is not isolated and should not be stated as one.
03cost
4K is a pixel count, not a quality tier
2.0 4K carries 0.025 bits per pixel against 720p's 0.167 — nine times the pixels and only 35% more data to describe them, at 4.9× the price. Normalised back to 720p it measures less detail than the 1080p file. Buy 4K only when you need the dimensions for a reframe.
04prompting
Dense segmented prompts provoke cuts rather than permit them
ARCANE asked for five hard cuts and both models produced eight detected discontinuities. IDOL asked for four and got ten on 2.5, fourteen on 2.0. Segment counts read as an invitation, not a ceiling — and 2.0 cuts harder than 2.5 on every advanced track. Budget for it, or cut in post.
05references
Bind references with an exclusion clause
A character sheet carries its pose, its framing, its flat studio light and its backing card along with the face you actually want. The LOCK prompt names all four as exclusions explicitly. Without that clause the reference tends to import its own lighting setup into your scene.
06grading
Flat near-black is the hardest thing you can ask for
Nothing posterised on the live-action tracks — shadow gradation held at 62–65 distinct levels through a hard grade. The one field that broke was the flat title card. Check shadow levels and not just blocking: Gemini Omni Flash did not block, it banded, at 48 levels where everything else sat between 57 and 65.
07casting
Three or more characters is the identity-drift cliff
SYNC puts four dancers on the same beat and IDOL puts five on a stage deliberately, because that is where reference ceilings start to matter — 2.5 allows 30 image references against 2.0's 9. Count your cast per segment when reviewing; a sixth body appearing is the classic failure.
08limits
Duration ceilings are hard walls
2.5 tops out at 30 seconds, 2.0 at 15. Past fifteen seconds 2.0 cannot compete at any price — it simply cannot go there. Write the RIG line as "model duration ceiling" rather than a number when you want a matched pair to stay byte-identical across two models with different limits.
09caveat
Cut counts here are detected, not verified
Every cut number on this page is ffmpeg scene detection at threshold 0.35. Whip pans, LED flashes and speed ramps all trip the detector, so these are stated as detected discontinuities rather than confirmed hard cuts. Treat them as a signal to go and look, not as a count.
10budget
Preflight the cost before every submission
Costs were checked with get_cost before each call and the closing balance matched the sum of per-clip costs to the credit. It is also how the "high bitrate is free" finding was confirmed before spending anything on it. Thirty-one renders came in at 1823 credits against a 1500 budget — the overage is published, not hidden.
11prompting
Negate sound, type and layout — never negate a shape
Naming a thing summons it, so a photoreal prompt should assert the positive: "every light source renders as a soft contained glow" beats "no lens flares". The exception is anything the model cannot draw. It cannot draw a drum loop, a subtitle or a panel gutter, so "no music, no captions, no split screen" is both safe and necessary. REDLINE ships a six-line block of exactly those negations and came back with clean sound-design-only audio and no grid artefacts, while every visual constraint in the same prompt is written as a positive.
12storyboard
Panel count sets beat count
A 15-second clip against a 6-panel board gives roughly 2.5s per beat and the model coasts between them — it renders a slideshow of beautiful frames. Against a 12-panel board it is ~1.25s per beat and it never gets the chance to settle. Rewriting the video prompt does not fix a slideshow; the cause is upstream and invisible if you only read the prompt. Judge the board with numbers first: per-panel mean luma and warm-pixel percentage take ten seconds to compute and catch the two failures that matter — an arc that does not arc, and an accent that does not carry.
06 — the takeaway
What to actually use
Always bitrate_mode: "high". Free on both models. No exceptions.
Photoreal action and budget coverage → 2.0 at 720p or 1080p. More detail per credit, and 31% cheaper.
Delivery master → 2.0 at 1080p, not 4K. Best bits-per-pixel on the ladder.
Anime, cel, any non-photoreal style → 2.5. Nearly twice the line energy.
Motion graphics and flat colour fields → 2.0 with high bitrate is the best file in the test. At default it is the worst.
Anything longer than 15 seconds → 2.5. Nothing else in the family goes there.
Large cast to hold consistent → 2.5 on paper: 30 image references against 9. Verify on SYNC before betting a job on it.
4K → only when you need the pixel dimensions for a reframe. Never for image quality.