Three AI video models arrived within twenty five days of each other and quietly rewrote what a single prompt can produce. Here is what each one actually does, what the benchmarks say, and what every second of output costs.
For most of the short history of AI video, one limit defined the medium: clips lasted a handful of seconds. A product ad needed stitching. A story needed cuts. A complete scene, with a beginning, a middle, and an end, simply did not fit inside one generation.
Between July 31 and August 24, 2026, that limit fell. ByteDance shipped Seedance 2.5 with 30-second single-pass generation and synchronized audio. MiniMax announced H3 the very same day, then published its weights, making a frontier-grade multimodal video model downloadable. Alibaba followed with Wan 3.0, matching the 30-second ceiling, adding document-to-video generation, and debuting at the top of the most watched community leaderboard.
This guide covers all three models in depth: the confirmed specifications, the honest caveats, the current benchmark standings, and verified pricing as of August 26, 2026. Every claim below traces back to launch documentation, official rate cards, or independent reporting.
Rather than lining up marketing claims, each model was assessed through five practical lenses, together called the DRIVE framework:
Sources include official announcements from ByteDance, MiniMax, and Alibaba Cloud, published rate cards from BytePlus ModelArk and the MiniMax platform, independent reporting from TechNode and The Decoder, and blind-vote rankings from the Artificial Analysis Video Arena. Where a number is provider-dependent or still unconfirmed, the text says so plainly.
Seedance 2.5 is ByteDance's flagship video model, previewed at the Volcano Engine conference in late June and released on July 31, 2026, with a public developer API following on August 7. It reached creators first through Jimeng AI and the Pro tier of Doubao, with API access through Volcano Engine Ark in China and BytePlus ModelArk internationally. It is also available online through platforms such as Seedance 2.5 generator.
The headline capability is a coherent 4 to 30 second clip generated in one pass, with dialogue, music, and sound effects produced together with the visuals rather than bolted on afterward. Multi-turn extension lets strong clips continue beyond that window while preserving lighting, action, and context.
Some creator products market 4K within the Seedance family, but independently verified API routes currently bill at 480p, 720p, and 1080p. Teams budgeting an API pipeline should treat 4K as a creator-product tier rather than a guaranteed endpoint. Seedance 2.5 also remains fully closed: there are no weights to download and no announced plans for a local version. Hands-on reviews describe the jump from Seedance 2.0 as a refinement in motion handling and prompt adherence rather than a dramatic quality leap, with the longer duration and expanded references doing most of the heavy lifting.
MiniMax H3 is the model behind the Hailuo line, announced on July 31, 2026 and published as open weights on Hugging Face on August 3. It is a 33-billion-parameter omni-modal system: text, images, video, and audio all enter one shared context, and the model produces video with native stereo audio, up to 15 seconds long, at up to 2K resolution and 24 frames per second. It can be tried online through hosts such as Topview's MiniMax H3 page as well as fal.ai and the Hailuo app.
The open release covers the H3-Base checkpoints (one for text and frame-guided generation, one for reference-based generation), while the Context-IR preprocessing layer that powers the full hosted pipeline remains closed. More importantly, the MiniMax H3 Community License excludes local deployment in the United States, the European Union, the United Kingdom, and South Korea unless MiniMax grants separate authorization, and it forbids training smaller models on H3's outputs. In practice, H3 is best described as regionally available open weights rather than a fully unrestricted release. Teams in excluded regions can still use hosted APIs.
On quality, the community verdict has been emphatic: H3 leads every open-weight bracket of the Artificial Analysis arena, ranks near the very top of the overall text-to-video board, and holds the top position in audio-inclusive video editing. For an openly downloadable model, that is unprecedented territory.
Wan 3.0 comes from Alibaba's Tongyi Lab. A public beta opened on August 6, 2026 through Alibaba Cloud Model Studio and Qwen Cloud, and the general release landed on August 24, one day after Alibaba closed a share placement of roughly 10.2 billion US dollars earmarked entirely for AI. Access runs through Alibaba's hosted services and partner platforms.
Like Seedance 2.5, Wan 3.0 generates up to 30 seconds in a single pass, double the 15-second ceiling of its predecessor Wan 2.7, at resolutions up to 1080p with audio generated by default. Alibaba reports improved instruction following, better consistency across shots, and higher audio quality, and cites real deployments across short drama, advertising, tourism promotion, and music videos.
Wan 3.0's signature capability is source-material input. A single request can include not just text, images, video, and audio, but also documents in DOC, XLS, PPT, PDF, TXT, Keynote, Pages, Numbers, and Markdown formats, plus public webpage URLs. The model reads the content and builds a video sequence from it, turning a pitch deck, a report, or a product page directly into motion. Among current mainstream video models, this document-to-video and webpage-to-video path is unique.
Wan 3.0 is closed. Alibaba built its video reputation on open releases, but the last open-weight Wan flagship remains Wan 2.2 from July 2025, and no Wan 3.0 files exist on Hugging Face, GitHub, or ModelScope. Any site advertising a Wan 3.0 download is not offering the real model. The hosted API also remains in preview, with approval required and endpoints documented in Beijing, Singapore, and Virginia.
The table below condenses the verified specifications of all three models as of August 26, 2026.
| Feature | Seedance 2.5 | MiniMax H3 | Wan 3.0 |
| Developer | ByteDance (Seed team) | MiniMax (Hailuo line) | Alibaba (Tongyi Lab) |
| Release | July 31, 2026 (API Aug 7) | July 31, 2026 (weights Aug 3) | Beta Aug 6; general Aug 24, 2026 |
| Max clip, single pass | 30 seconds (4 to 30s range) | 15 seconds | 30 seconds |
| Max output resolution | 1080p on verified API routes; 4K marketed in creator products | 2K native (1440p short edge), 24fps | 1080p |
| Native audio | Yes, video and audio in one pass | Yes, native stereo audio | Yes, generated by default |
| Inputs | Text, image, video, audio | Text, image, video, audio | Text, image, video, audio, documents, webpages |
| Reference capacity | Up to 50 assets: 30 images, 10 videos, 10 audio | Hosted endpoints: up to 9 images, 3 videos, 3 audio | Omni-Reference for characters, products, and sets |
| Editing controls | Second-level timestamps, storyboards, local segment editing, extension | Instruction-led editing, V2V motion transfer, first and last frame | Smart duration, extension, cross-shot consistency |
| Open weights | No | Yes, with regional license limits | No (Wan 2.2 remains the open flagship) |
| First-party pricing | About $0.10/s at 480p and $0.23/s at 720p (token-billed) | $0.13/s at 2K; roughly $0.08/s at 768p | $0.05 / $0.10 / $0.20 per second at 480p / 720p / 1080p |
| Primary access | Jimeng AI, Doubao Pro, Volcano Engine Ark, BytePlus ModelArk | Hailuo app, MiniMax API, Hugging Face, third-party hosts | Model Studio, Qwen Cloud, hosted API only |
Table: Confirmed specifications from launch documentation and provider rate cards, checked August 26, 2026.
Clip length is the clearest single number separating this generation from the last. A 30-second window covers a complete broadcast spot, a full product story, or a scene with an actual arc, all without stitching. Seedance 2.5 and Wan 3.0 both reach it in one pass. MiniMax H3 stops at 15 seconds, though its 2K output resolution is the highest of the three.
_1788172246.jpg)
Figure 1: Maximum single-pass clip length. Gray bars show the previous generation and a leading rival for context.
The most credible public quality signal in AI video is the Artificial Analysis Video Arena, where users vote between two unlabeled clips generated from the same prompt. As of late August 2026, Wan 3.0 leads the text-to-video board for models with audio at an Elo of 1241, narrowly ahead of Google's Gemini Omni Flash at 1237 and MiniMax H3 at 1226.
_1788172238.jpg)
Figure 2: Artificial Analysis text-to-video arena, models with audio, late August 2026.
Three caveats keep this honest. First, Seedance 2.5 does not yet appear on the public board, so its predecessor Seedance 2.0 stands in for the family; its own ranking is still forming. Second, gaps of a few Elo points sit within normal voting noise, so the top three are better read as a cluster than a podium. Third, image-to-video tells a different story: there, the Seedance family holds the top audio-inclusive spot, with H3 close behind. Rankings on these boards move week to week, so a live check is always worth thirty seconds before a big production decision.
Published first-party rates, checked on August 26, 2026, are the fairest baseline, since third-party hosts add their own margins and the same generation can cost anywhere from ten cents to over a dollar per second depending on the route.
_1788172258.jpg)
Figure 3: First-party list rates per second of generated video. Seedance 2.5 figures are derived from BytePlus ModelArk's published 5-second examples.
In practical terms: a full 30-second 1080p Wan 3.0 clip lists at $6.00, or $4.20 during the launch promotion. A 30-second 720p Seedance 2.5 clip lands near $6.94. H3's 15-second maximum at 2K costs $1.95. The real budget driver, though, is rerolls: a model that needs five attempts per accepted clip is more expensive than its rate card suggests, so cost per accepted output is the number worth tracking.
Each model lives on a different first-party platform, which makes side-by-side testing tedious. Multi-model platforms solve this: Topview, for example, runs workflows for Seedance 2.5, MiniMax H3, and Wan 3.0 under one roof, alongside storyboard tools, previs, and conversational editing, so the same prompt and references can be tested across models before a campaign locks in. Whichever route a team picks, the sensible first step is the same: run one identical brief through all three and judge the results against the actual deliverable, not the demo reel.
Be the first to post comment!