Seedance 2.5 — what ByteDance's new video model can do

Seedance 2.5 — what ByteDance's new video model can do

On July 31 ByteDance released Seedance 2.5 — the next generation of its flagship video model. For now the model is available in the company's Chinese products — the Jimeng AI web app and the Pro tier of Doubao — with the API on Volcano Engine Ark opening in early August. As soon as the model is available via API, we will add it to GPTunneL. Here is what we know so far, with the official demos.

Release timeline

The model was announced on June 23 at the Volcano Engine FORCE conference in Beijing, then still in enterprise beta. A limited public test ran in early July, and the consumer launch followed on July 31 — China first. Alongside it, Seedance 2.0 gained 4K output.

The headline: one-take video up to 30 seconds, with sound

The key difference from Seedance 2.0 is duration. The model generates a continuous clip of up to 30 seconds in a single pass, with no stitching of short fragments — and the audio is generated together with the video in that same pass, not layered on top of a finished clip. For comparison, most competitors top out at 8–15 seconds per generation.

The flagship demo is a 30-second one-take: a character in a black coat walks through six connected rooms, each shifting the style and mood of the shot. Another example is a full music video with a single performer, built from one prompt using appearance and wardrobe references — no timeline, no keyframes. And to show the full pipeline, ByteDance released "The Missing Pair", a short film made entirely with Seedance 2.5.

Multi-round extension grows a clip while keeping characters, environments and pacing, and a beta Ultra-Long Video mode spotted in the Jimeng app goes up to 180 seconds.

Here is what that looks like in the official demos. A dance in front of a wall of fire — one continuous shot with music and motion generated in a single pass:

And a 30-second mini-story of a suitcase travelling through an airport — camera, light and "cast" keep changing while the object stays consistent the whole way:

Up to 50 references per generation

Seedance 2.5 accepts up to 50 multimodal inputs: 30 images, 10 videos and 10 audio tracks. Two stand-outs:

  • clay-render references (white-model control) — untextured 3D blocking of a scene that pins down poses, motion paths and camera angles; per ByteDance, the model then builds physically correct light on its own — direction, color temperature, shadows;
  • audio as a reference — the model can follow an uploaded soundtrack, matching pacing and motion to it.

According to Artificial Analysis, Seedance 2.0 already led the image-to-video leaderboard among models with audio — 2.5 builds on that base.

Editing without reshoots

The second big theme of the release is editing finished clips:

  • timestamp edits — "replace the background from second 12 to 18";
  • green-screen background replacement that adapts the subject's physics and lighting;
  • re-shooting the camera move on an existing clip;
  • local edits of a frame region from a reference — without regenerating the whole video.

Realism also takes a visible step up: textures, skin, light and facial acting look noticeably richer than in 2.0. Judge for yourself — a barbershop scene with two characters and a mirror (mirrors are a classic failure mode for video models):

To its credit, ByteDance admits that complex motion and scenes with several closely interacting subjects are still hard for the model.

How Seedance 2.5 stacks up

  • Sora 2 is effectively out of the race: OpenAI shut down the Sora app and web version on April 26, and the API goes offline on September 24, 2026.
  • Kling 3.0 — native 4K at 60 fps, but up to 15 seconds per pass.
  • Veo 3.1 — the best synchronized audio with dialogue, but 8 seconds per generation.

Seedance 2.5's strengths are duration, reference-driven control and reshoot-free editing. This is a bid for a production workhorse, not just "the prettiest clip". The scale it can handle is impressive too:

API timing and pricing

Volcano Engine has published its rates: roughly ¥70 per million tokens for generation without video input and ¥42 with video input; on the international BytePlus platform — $10.70 and $6.40 respectively. In practical terms that is about $0.70–1.80 per 30-second clip at base resolution; 4K will cost several times more. Exact GPTunneL pricing will arrive together with the model.

While we wait — try Seedance 2

Seedance 2 is already available in GPTunneL — here is a taste of what it can do:

Generate right in the lab — no setup, local payment methods supported. And once Seedance 2.5 is available via API, the model will land here: follow the news in our blog and Telegram channel.