LTX-2.5 vs MiniMax-H3

Ten LTX-2.5 clips generated 2026-08-17 on one B200, against the H3 clips already on disk. Same 768×1344 at 24fps, same in-model audio, same concepts and verbatim dialogue. Clips are full quality — nothing re-encoded — so skin texture is judgeable. Unmute to check the audio; both models generate it in-model.

~$0.13LTX per 15s clip
~$0.97H3 per 15s clip
10 / 10LTX clips succeeded
66smean LTX gen time

What to look for. The result splits by conditioning path, and that matters more than the price.

Image-conditioned (green tag): LTX is genuinely competitive — identity, set, wardrobe and props hold for the whole clip. H3 still runs cooler and more matte, which is the look the house prompts were tuned for, and it holds framing where LTX wanders.

Text-only (red tag): LTX drifts. In walk_p1 the person changes and the lighting jumps mid-shot. H3's ref2va holds a presenter across scene changes and LTX-distilled has no equivalent — that gap, not the cost, is what decides adoption for recurring-cast work.

grocery_trailer_voice

image-conditioned

Best case for LTX. Identity, kitchen, wardrobe and props hold for the whole clip, and the comedy arc lands: amused → wide-eyed → the mock-trailer bellow → collapsing into laughter.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 243f · 10.12s · 32000Hz · 1.8MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 241f · 10.04s · 48000Hz · 3.4MB

part1_hook

image-conditioned

15s image-conditioned. Identity holds across all 361 frames. LTX pulls in closer and is more animated; H3 holds the reference framing more strictly.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 362f · 15.08s · 32000Hz · 2.1MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 5.0MB

smoke_short

image-conditioned

Short control clip. Note LTX drifts off-prompt near the end — stops addressing camera and drops the earbud the prompt named.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 107f · 4.46s · 32000Hz · 0.6MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 105f · 4.38s · 48000Hz · 2.1MB

walk_p1

text-only

THE FAILURE CASE. Watch the LTX side: the person visibly changes partway through, and the lighting jumps from cool daylight to hard golden sunset mid-shot. H3's ref2va holds one presenter and one lighting state for 15s.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 362f · 15.08s · 32000Hz · 4.2MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 15.8MB

part2_reveal

text-only

Text-only bedroom. H3 uses ref2va to keep the same presenter in a new room; LTX has no equivalent, so this is a fresh person each time.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 362f · 15.08s · 32000Hz · 2.1MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 4.2MB

part3_payoff

text-only

Text-only bathroom. Same structural gap as above.

MiniMax-H34×B200 · 24 steps · ~$0.97
768×1344 · 362f · 15.08s · 32000Hz · 2.3MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 9.5MB

bedroom_morning_show

text-only

Incidental win for LTX: it came out correctly PORTRAIT. H3's only version of this concept is 1344×768 LANDSCAPE — the ref2va + aspect_ratio:auto trap, a full-price unusable clip.

MiniMax-H34×B200 · 24 steps · ~$0.97
1344×768 · 192f · 8.0s · 32000Hz · 1.0MB
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 193f · 8.04s · 48000Hz · 2.1MB

hook_kitchen_02

image-conditioned

No H3 counterpart — never generated.

No H3 counterpart
never generated
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 7.0MB

hook_commute_01

text-only

No H3 counterpart — never generated.

No H3 counterpart
never generated
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 14.0MB

hook_desk_03

text-only

No H3 counterpart — never generated.

No H3 counterpart
never generated
LTX-2.51×B200 · 8+4 steps · ~$0.13
768×1344 · 361f · 15.04s · 48000Hz · 4.1MB

Character consistency — LTX-2.3 + Ingredients IC-LoRA

The only character-consistency adapter in the LTX family. Left is LTX-2.5 text-only on the same concept; right is Ingredients. It works — but it is a 5-second landscape model on the older 2.3 base.

walk_p1_id

reference sheet

The headline pair. LTX-2.5 (left) changes the person and jumps the lighting partway through 15s. Ingredients (right) holds one person and one lighting state. NOTE it is only 5s — a third the duration is a third the chance to drift, so this shows the mechanism works, not that it survives 15s.

LTX-2.5 text-only1×B200 · drifts mid-clip
768×1344 · 361f · 15.04s · 48000Hz · 15.8MB
LTX-2.3 + Ingredients768×448 · 121f · in distribution
768×448 · 121f · 5.04s · 1.1MB

part2_reveal_id

reference sheet

Bedroom scene. Same presenter carried from the reference sheet into a room that is not in the sheet.

LTX-2.5 text-only1×B200 · drifts mid-clip
768×1344 · 361f · 15.04s · 48000Hz · 4.2MB
LTX-2.3 + Ingredients768×448 · 121f · in distribution
768×448 · 121f · 5.04s · 0.4MB

part3_payoff_id

reference sheet

Bathroom scene. Note the softness — 0.34 MP against LTX-2.5's 2.09 MP at 1080p.

LTX-2.5 text-only1×B200 · drifts mid-clip
768×1344 · 361f · 15.04s · 48000Hz · 9.5MB
LTX-2.3 + Ingredients768×448 · 121f · in distribution
768×448 · 121f · 5.04s · 0.4MB

walk_p1_ood

out of distribution

Forced to OUR house format: 9:16, 15s. The card calls this out of distribution and it shows. It runs without erroring — it degrades rather than fails — but this is why Ingredients is not adoptable as-is.

LTX-2.5 text-only1×B200 · drifts mid-clip
768×1344 · 361f · 15.04s · 48000Hz · 15.8MB
LTX-2.3 + Ingredients768×1344 · 361f · OUT OF DISTRIBUTION
768×1344 · 361f · 15.04s · 1.9MB

Audio-to-video — driven by real Lissin TTS

The audio is your own Gemini-3.1-flash TTS from the voice feed, passed through unmodified — the clip carries the actual Lissin voice, not a re-synthesis. Video prompts were written per clip to match each voice's director prompt and timbre. Left is with NO image conditioning, right conditions frame 0 on a photoreal gpt-image-2 still. Unmute these.

sports_hype

audio-driven

sports-01 · Play-by-Play Hype · voice Fenrir · excitable
Video built to peak where the audio peaks. Tests whether A2V reads energy out of the waveform.

no conditioningimages=[] — pass 1
1088×1920 · 385f · 18.82s · 48000Hz · 27.3MB
gpt-image-2 frame 0photoreal still conditions frame 0
1088×1920 · 385f · 18.82s · 48000Hz · 10.1MB

noir_detective

audio-driven

true-crime-01 · Noir Detective · voice Algenib · gravelly
Deliberately near-static staging so the audio's long pauses have somewhere to land.

no conditioningimages=[] — pass 1
1088×1920 · 553f · 25.9s · 48000Hz · 9.9MB
gpt-image-2 frame 0photoreal still conditions frame 0
1088×1920 · 553f · 25.9s · 48000Hz · 13.2MB

sleep_meditation

audio-driven

wellness-01 · Sleep Meditation · voice Vindemiatrix · gentle
Total camera stillness is the point — motion here would mean A2V is not reading the audio's pace.

no conditioningimages=[] — pass 1
1088×1920 · 649f · 29.98s · 48000Hz · 62.2MB
gpt-image-2 frame 0photoreal still conditions frame 0
1088×1920 · 649f · 29.98s · 48000Hz · 46.4MB