Every few months a video model arrives with a per-second price that looks impossible next to its neighbours, and the reaction is always the same: it must be worse somehow. Lower resolution, shorter clips, some hidden throttle. Occasionally that’s true. More often the price gap is real and the confusion comes from comparing the wrong numbers.
Per-second generation cost is the number everyone quotes and close to the least useful one available. What a team actually spends is determined by three multipliers sitting on top of it, and those multipliers vary between models far more than the headline rate does.
Multiplier one: what’s included in a second
The first distortion is that “one second of video” means different things depending on the model.
Most video models output silent picture. To get a usable second you then need a voice model, a sound library or a sound designer, and time spent aligning the two. Those costs are real; they’re just accounted for somewhere else in the budget, which makes the generation invoice look flattering.
Minimax H3 generates native stereo audio in the same pass as the picture — dialogue, ambience, foley, and music, with their sync relationships already intact. Whatever the rate card says, that second arrives finished. Comparing it against a silent second from another model is comparing a delivered shot to a raw plate.
This alone accounts for a meaningful slice of the apparent gap, and it’s the slice most comparisons ignore entirely because it doesn’t appear on anyone’s pricing page.
Multiplier two: attempts per usable clip
Here’s the number that actually governs spend: how many generations does it take to get one you can use?
Nobody’s first render is right. The real unit cost is the sticker price times the attempt count, and attempt count is where workflows diverge sharply. Under a generate-only model, fixing anything means regenerating the whole clip — and regeneration changes everything, not just the thing you wanted changed. The lighting shifts, the face drifts, the approved parts stop being approved. So you re-roll, and re-roll, and the effective cost per usable second climbs to some multiple of the advertised one.
H3 edits locally instead. Replace an object, change a background, adjust a performance, rewrite a line and have the mouth follow the new words — while the rest of the frame stays where it was. It sits first on the Artificial Analysis video editing leaderboard, ahead of Seedance 2.0, which matters here for a purely financial reason: a targeted edit is one operation, while a re-roll is a lottery ticket that also discards work you’d already paid for.
If model A costs twice as much per second as model B but needs a third as many attempts to converge, model A is cheaper. Most published comparisons never get as far as asking.
Multiplier three: variants
The third multiplier is the one that turns a small gap into a large one, and it’s specific to commercial work.
Nobody delivers one video any more. A single brief produces a 16:9 master, a 9:16 cut, a square version, a short pre-roll, several hook variants for testing, and language versions per market. Every one of those is a generation, or several.
At one clip, a per-second difference is a rounding error. At forty clips a month across an agency’s accounts, it’s a line item someone has to defend. The full rate structure is on the Minimax h3 pricing page, and the useful exercise is to multiply it by your genuine monthly variant count rather than by one hero spot — that’s the number that decides whether the difference is interesting or decisive.
The spec discipline nobody mentions
Part of the answer is also that H3 doesn’t sell specs it doesn’t need.
Output runs 5–15 seconds at 24 FPS, up to 1440p. Those aren’t maximal figures, they’re delivery figures. 24 FPS is the frame rate of film and television, so output conforms into a real timeline without conversion. 1440p is the resolution at which typography, product detail, and interface elements stay legible through motion — which is the threshold that actually matters for commercial work, as opposed to a 4K number that mostly gets downscaled before anyone sees it.
Generating pixels and frames that get thrown away in the edit is a cost with no corresponding benefit. A model targeted at what gets delivered rather than at what benchmarks well can price accordingly. Teams comparing Minimax H3 AI Video against higher-spec alternatives usually find the deciding question isn’t whether 4K would be nice, but whether they’d ever have shipped it.
What can’t be answered honestly
It’s worth being straight about the limits of this explanation.
The parts above are structural and observable: what’s bundled into an output, how many attempts a workflow requires, how many variants a brief generates, which specs are being targeted. Anyone can verify them against their own usage.
What nobody outside MiniMax can tell you is the inference-level reason — architecture, model size, serving efficiency, or how much of the current rate reflects genuine cost versus a launch position in a competitive market. Vendors don’t publish that, and articles claiming to explain it are guessing. Prices in this category have also moved fast in both directions, so anything quoted today is a snapshot rather than a structural fact.
How to actually compare
Ignore the per-second headline for a moment and compute cost per usable finished second. Take the rate, multiply by attempts to convergence on your own material, add whatever audio work each option requires afterwards, then multiply by the number of variants a typical brief produces.
Run that on a real job rather than a demo prompt. The ranking it produces is frequently different from the one on the pricing pages — and it’s the only ranking that shows up in a budget



