The AI video generation market shifted decisively in 2026. Sora (OpenAI) was permanently shut down on April 26, 2026. Three tools now dominate the field: Seedance 2.0 (ByteDance), Kling 3.0 (Kuaishou) and Runway Gen-4.5. For creative studios, the question is no longer "should we use AI video?" but "which tool for which use case?". Here is our comparison after 18 months of production use at Nexia.
Seedance 2.0 is currently the highest-performing tool on the market - it outranks Runway, Kling and Veo on Elo benchmarks. Its strength: native audio-visual generation (dialogue, SFX, music) in a single pass, up to 2K resolution, with up to 9 reference images. It is the tool we use when the brief demands a premium cinematic output with audio synchronisation. Points to note: only accessible via CapCut/Dreamina for now, API planned for Q3 2026. Ideal for brand films and emotional content with narration.
Kling 3.0 offers the best quality-to-control ratio on the market. Its multi-shot storyboard mode lets you direct the sequence flow - essential for motion designers who want narrative consistency, not surprises. Clips up to 15 seconds, 4K output, integrated audio generation. Limitations: artefacts still present on complex facial expressions. Ideal for motion design, product content and character-driven sequences.
Runway Gen-4.5 remains the Western reference for camera control. Motion Brush 2.0 and Camera Director enable precise camera movements - pan, zoom, tracking. Quality is slightly below Seedance 2.0 but the level of control is unmatched for compositing and assembly work. Limitations: slower than Seedance, weaker on native audio. Ideal for productions requiring precise camera movements and a compositing workflow.
The tool does not make the film - the art director makes the film. The tool accelerates.
In Nexia production 2026, we combine tools according to the brief. Seedance 2.0 for premium final deliverables with audio. Kling 3.0 for creative control in motion design. Runway Gen-4.5 for camera movements and compositing. Luma Dream Machine 3.5 remains in use for rapid prototyping during creative exploration.
The question is no longer "should we use AI video?" but "which tool for which use case?"
We rewrite this page when our own production changes, not when a vendor publishes a press release. Since April, three things moved in our pipeline, and all three cost us money to find out.
1. Our default engine is now Google Flow. Not because it wins a benchmark, but because of the ratio nobody publishes: matter per credit. On our account, image-to-video costs 12 credits for 8 seconds and 15 credits for 10 seconds, with up to 7 reference images per shot. The lighter Veo 3.1 model runs 8 seconds for 10 credits. Ten seconds at 15 credits is the best value we have measured, and it is what an editor can actually cut. On 31 July we produced 15 clips this way and cut a full trailer out of them.
2. Credits expire, and that belongs in the price. A monthly allowance that does not roll over is not the same product as a pay-as-you-go balance, even at the same headline price. Ours lapsed in July and the engine we had built a whole pipeline around went dark overnight. When you compare tools, compare the billing model with the same seriousness as the output quality: a studio plans a month of production, not a demo.
3. Open-weight models are now a real third channel, with a real cost. A single 12 GB consumer GPU (RTX 3080 Ti) runs a library of open video models locally. For image-to-video quality, Wan 2.2 is the one worth the disk space. Two honest numbers before anyone calls it free: a 5 second shot at 832x480 took over 30 minutes on that card, and each model family weighs 15 to 60 GB before the first pixel. What it buys is not speed, it is confidentiality: nothing leaves the machine, which is the only thing that matters on an NDA project.
Local models ship in several numeric formats, and the format is not a quality setting: it is a hardware requirement. FP8 weights only have hardware support from RTX 40 series onwards. On an RTX 30 the maths is emulated, and the failure is deceptive: the first frame is perfect, because it comes from your source image, and the video degrades into mush around frame 40. We downloaded 27 GB and spent 37 minutes generating before seeing it.
The rule we now apply before any download: read the weight format, then check it against the GPU generation. FP8 needs RTX 40 or newer. FP4 and NVFP4 need RTX 50. BF16 and int8 run everywhere. It is a thirty second check against tens of gigabytes.
One more reading habit, in the same spirit: when a model catalogue says "recommended for cinematic video with native audio", that column describes features, not a quality ranking. The model recommended there may simply be the only one that generates a synchronised soundtrack, which is precisely the thing we mute on every single delivery.
Every quarter a model announces longer shots in a single pass. The number that matters is not the maximum length, it is the length that survives an edit. We never generate under 15 seconds on action shots, because under that there is nothing to cut. A demo reel proves a model can hold 20 seconds; a timeline proves whether second 14 still holds the character, the light and the camera.
So this page will keep saying the same thing in different words: the ranking that counts is the one you measure on your own brief, in the month you deliver it.
A benchmark tells you what a model can do once. Production tells you what it does sixty times in a row.