Vogoo.ai logoVogoo.ai
  • Prompts
  • Blog
  • Pricing
Vogoo.ai
Audio to Video AI: Seedance 2.5's Audio-Only Mode
2026/08/06

Audio to Video AI: Seedance 2.5's Audio-Only Mode

Seedance 2.5 adds audio-only reference: drive pacing, beat cuts, and lip sync from one music or voice clip. Audio to video AI guide with official demos.

Every AI video workflow has the same missing step: you generate a silent clip, then spend an hour in an editor trying to make the cuts land on the beat and the mouths match the words. Seedance 2.5 attacks that problem from the other direction — it lets you hand the model a piece of audio first, as a reference input, and generate picture that follows the sound. This is the new audio-only reference mode: feed in a BGM track, a voice line, or a sound effect, and the model uses it to drive pacing, beat sync, and lip alignment from the very first frame. It's a genuine first for the Seedance line, and as of this writing there's almost no tutorial coverage of it anywhere.

This guide is based on ByteDance's official Seedance 2.5 enterprise practice guide, cross-checked hands-on in Vogoo's studio. Every spec, limit, and prompt below traces to the official documentation or a source listed at the end — nothing is guessed.

What is audio-only reference?

Seedance 2.5 is what ByteDance calls an audio-video joint generation model: sound and picture are generated together, not stitched in post. Its reference system accepts three modalities — images, video clips, and audio clips — and 2.5 is the first version where audio can stand alone as the only reference.

The official practice guide lists exactly seven supported modality combinations:

#CombinationStatus in 2.5
1Images onlyCarried over
2Video onlyCarried over
3Audio onlyNew in 2.5
4Images + videoCarried over
5Audio + videoCarried over
6Images + audioCarried over
7Images + audio + videoCarried over

Per the guide, an audio reference can be a BGM track, a voice line, or a sound effect — the model reads rhythm, mood, and speech from the clip and builds picture to match: cuts land on beats, character mouths align to spoken lines, and scene energy follows the track's dynamics.

The audio upgrade sits inside a much larger reference expansion from 2.0 to 2.5:

SpecSeedance 2.0Seedance 2.5
Single-shot video length15s30s
Reference images930
Reference video clips310
Reference audio clips310
Total reference duration15s30s (video and audio counted separately)
Audio as sole referenceNot supportedSupported
Native languages—11 (Chinese, English, Spanish, Japanese, Korean, Arabic, and more)

If you want the broader picture of what changed beyond audio, the Seedance 2.5 hub covers the full release; this article stays focused on the audio pipeline.

Audio reference specs: what you can upload

Before building prompts, know the hard limits. These come straight from the official practice guide:

ConstraintLimit
Number of audio clips0 to 10 per generation
File sizeUp to 15MB per clip
Formatsmp3, wav
Single clip duration2 to 30 seconds
Total audio durationUp to 30 seconds across all clips
Interaction with video referencesAudio and video durations are counted separately — you can upload 30s of video references and 30s of audio references in the same job

Two practical notes. First, avoid the extremes: a clip near the 2-second floor carries too little information for the model to lock onto, while one that fills the whole 30-second ceiling can dilute the key features you want referenced. A clean excerpt of just the section you care about works better than the full track. Second, the @-syntax discipline applies to audio exactly as it does to images: every referenced clip needs a usage statement immediately after it — "sound design follows @Audio 1" — never a bare, floating "@Audio 1" with no stated role.

Five official cases, from pure audio to full multimodal

The practice guide demonstrates the audio pipeline across a ladder of complexity. Here are the five official demos, each with its real prompt where the guide publishes one.

Case 1 — Audio only: two geckos plan a jailbreak

This is the flagship demo of the new mode. The only references are two audio clips — a circus-backstage ambience bed and a music track. No images, no video. The model builds a 3D-animated scene whose staging and emotional beats follow the sound.

Official Seedance 2.5 demo — ByteDance / Volcano Engine practice guide; prompt translated from the original.

The official prompt opens like this:

[Two Geckos' Jailbreak Meeting — 3D animation | circus backstage | ~20 seconds]

[One-line synopsis] Open with an ambient establishing shot that says "this is
the circus backstage." Then two geckos (one chubby, one tiny) plot their
escape from the circus, break into a brief argument, and finally reach an
agreement. Amplified backstage ambience carries the base layer, with
background music underpinning almost the whole runtime and swelling at key
emotional beats.

Notice what the prompt does: it explicitly assigns each audio layer a job (ambience = base layer, music = emotional emphasis). That's the pattern to copy.

Case 2 — Rhythm transfer: the headphone ad

This case pairs a product image with a reference video, and the target is the edit rhythm — the cut points and transition speed of the original ad get transferred to a brand-new product film. It's the same beat-matching muscle the audio mode uses, controlled from the video side:

Official Seedance 2.5 demo — ByteDance / Volcano Engine practice guide (prompt translated from the official guide).

Follow the camera work and editing rhythm of @Video 1 to generate a short ad
for the headphones in @Image 1. Match the scenes to the product's tone —
modern, clean, tech-forward. Keep the cut points between product close-ups
and wide shots identical to the original, with the same motion speed and
transition style.

Case 3 — Three modalities at once: the walking one-shot

The most complete workflow: a stack of reference images (character, wardrobe, scenes), reference videos (camera language), and a sound-effect clip, all in one 26-second one-shot with day-night and seasonal transitions. The opening of the official prompt:

Official Seedance 2.5 demo — ByteDance / Volcano Engine practice guide; prompt excerpt translated from the guide.

Core instruction: a 26-second one-shot narrative short film with stable
tracking, weaving the follow-cam of @Video 1 with the smooth orbiting camera
of @Video 2. A smooth sense of forward motion. Day-night alternation and
four-season transitions happen inside the single shot. The protagonist is a
European woman @Image 1, immersed in a lively sea of people, emphasizing
extreme solitude and cinematic photography.

0-3s (steady back-follow): the old wooden door @Image 2 creaks open, and the
camera follows close behind the European woman, dressed as in @Image 3, as
she steps out...

Every asset is named and given a role — timestamped, segment by segment.

Case 4 — Music-driven multi-character: the ballroom

A 41-second palace-party group scene with a large cast held consistent across the shot, where a reference audio clip controls a character's voice timbre for a spoken line:

Official Seedance 2.5 demo — ByteDance / Volcano Engine practice guide, with the prompt line translated from the original.

The key line from the official prompt:

Man 1 reaches in from the front-left of the frame, picks up a champagne glass
with his right hand, raises it in front of his right shoulder, and calls out
loudly to the guests ahead: "Everyone, enjoy this party!" (voice timbre
follows @Audio 1)

That parenthetical is the entire lip-sync workflow: script the line in the prompt, point the timbre at a reference clip, and the model performs it — mouth movement included.

Case 5 — Music over a 3D white model: the spaceship

The guide's hardest reference test: a professional 3D white-model (untextured previz) animation is fed in as @Video 1, a style frame as @Image 1, and the model re-renders the entire sequence at film quality while preserving structure and camera moves frame-for-frame. The official demo cut runs against a music track, showing how previz becomes a finished, scored sequence:

Official Seedance 2.5 demo — ByteDance / Volcano Engine practice guide; prompt rendered in English from the guide's original.

Keep the camera motion, duration, composition, framing, spatial relationships,
object positions, model structure, and motion paths in @Video 1 unchanged.
Use @Image 1 as the reference for materials, lighting, color, and overall
mood. Replace the white-model materials in @Video 1 with realistic materials
close to @Image 1, adding natural light and shadow, contact shadows, ambient
light, reflections, highlights, and spatial layering, for a true film-grade
render throughout.

You can try this same reference stack — audio, images, and video together — for free in Vogoo's Seedance 2.5 studio, which supports the model's full 30-second one-shot generation, up to 4K output, native synced audio, and up to 50 references per job.

Building your own audio + image / video prompts

Distilled from the official cases, here's the repeatable workflow:

  1. Trim the audio to the section that matters. mp3 or wav, under 15MB, and each clip must run between 2 and 30 seconds — a short, focused excerpt beats the full track.
  2. Upload and declare every clip's job. A canonical pattern for mood-driven generation: Music mood follows @Audio 1, visual subject follows @Image 1 — generate a short film that matches the rhythm. One asset, one declaration, always.
  3. Separate layers when you use multiple clips. Ambience bed, music, voice, and spot effects each get their own clip and their own stated role — exactly like the gecko case.
  4. Script spoken lines in the prompt, point timbre at audio. Write the exact dialogue in quotes and append (voice timbre follows @Audio N) — the ballroom pattern.
  5. For swapping sound on an existing video, use the editing pattern: Keep the picture and camera work of @Video 1 unchanged; replace the background music with @Audio 1. This is a video-to-video operation rather than a fresh generation.

If you're new to the @-reference grammar itself, the broader reference-to-video workflow uses the same syntax across all modalities.

What makes this different

Most AI video tools treat audio as an afterthought — you generate silent picture from a text or image prompt, then bolt sound on in an editor and hope the cuts roughly land. A few models generate their own soundtrack, but you can't hand them your track and have the visuals obey it. Seedance 2.5's audio-only reference inverts the pipeline: your audio is a first-class input that shapes rhythm, cuts, mood, and mouths at generation time. Open-weight rival MiniMax H3 pushed hard on motion quality this cycle, but reference audio as a standalone driving input remains the differentiator ByteDance is claiming — and per the enterprise guide, 2.5's ten-clip, 30-second audio budget is counted separately from the video-reference budget, so neither modality crowds out the other.

Where audio-driven generation fits

Three workflows get dramatically shorter:

  • Music video beat cuts. The guide's beat-sync cases show outfit and scene changes landing on the soundtrack's beats — the model cuts where the music tells it to. This is the core move behind any AI music video generator workflow: pick the track first, let the picture chase it.
  • Talking-head and dubbed content. Scripted lines plus timbre reference gives you AI lip sync from audio at generation time, and 2.5's native support for 11 languages means the same scene can be performed across markets with mouths matching each language.
  • Mood films and ambience pieces. The gecko case shows that a sound bed alone can carry scene-setting — useful for brand mood reels, game atmosphere pieces, and title sequences where the score exists before any footage does.

FAQ

What audio formats does Seedance 2.5 accept as reference?

mp3 and wav, up to 15MB per file. Each clip must run between 2 and 30 seconds, you can attach up to 10 clips, and total audio duration is capped at 30 seconds per generation.

Can Seedance 2.5 generate a video from only a song or audio clip?

Yes — that's the new audio-only reference mode in 2.5. You upload one or more audio clips with no images or video, describe the scene in text, and the model generates picture whose pacing, cuts, and mood follow the sound. Earlier Seedance versions required at least an image or video reference.

Does audio reference handle lip sync?

Yes. Write the exact dialogue in your prompt and attach a voice clip with a declaration like "voice timbre follows @Audio 1" — the official ballroom demo uses exactly this to make a character deliver a spoken line with matched mouth movement. The model also aligns lip sync across its 11 supported languages.

Do audio references reduce how much video I can reference?

No. The official guide states audio and video durations are counted separately — up to 30 seconds of video references and 30 seconds of audio references can coexist in one job, inside the overall 50-asset cap (30 images, 10 videos, 10 audio clips).

How is audio to video AI different from adding music in an editor?

Editing adds sound after the picture is fixed, so you're re-cutting footage to fit the track. Audio-driven generation makes the track an input: beats, energy shifts, and spoken lines shape the visuals as they're created, so cuts and mouths land right without manual sync work.

Where can I try Seedance 2.5's audio features?

The Seedance 2.5 studio on Vogoo runs the model in the browser — 30-second one-shot generation, up to 4K, native synced audio, and up to 50 reference assets per job, with a free way to start. Plan details live on the pricing page.

The bottom line

Audio-only reference is the quiet headline of Seedance 2.5: for the first time in this model line, a music track, a voice clip, or a sound effect can be the thing that drives a generation, with picture built to fit the sound instead of the other way around. The specs are generous — ten clips, 30 seconds of audio counted separately from video references — and the official cases show the same @-syntax working from a two-clip gecko short all the way up to a multi-character ballroom scene with timbre-matched dialogue. If your work starts from sound — a beat, a script, a score — this collapses the sync step that used to eat your afternoon. Trim a clip, declare its job, and run it in Vogoo's Seedance 2.5 studio to hear and see the result in one pass.

Sources

  • Doubao-Seedance-2.5 Enterprise Practice Guide — ByteDance / Volcano Engine (Lark wiki) — primary source: all audio specs (15MB, mp3/wav, 2–30s clips, 10-clip / 30s caps, separate duration accounting), the seven modality combinations, and every official case prompt quoted above.
  • Seedance 2.5 — One-take Creation, Flexible Referencing — ByteDance Seed blog — official announcement confirming 30 images + 10 video clips + 10 audio clips as references and 30-second single-pass audio-video generation.
  • Seedance 2.5 official product page — ByteDance Seed — positions 2.5 as an audio-video joint generation model; confirms 30s single generation, white-model control, and reference-based editing.
  • ByteDance launches Seedance 2.5 video-generation model — TechNode — July 31, 2026 launch on Jimeng AI and Doubao Pro, with API access planned via Volcano Engine Ark.
  • ByteDance unveils Seedance 2.5, a 30-second native 4K AI video model that accepts 50 reference inputs — TNW — independent confirmation of the 50-reference cap and native 4K output.
  • ByteDance Unveils Seedance 2.5 Video Model — The Information — industry-press briefing on the announcement.
  • ByteDance Seedance 2.5: Native 30-Second AI Video, No Stitching Required — Tech Times — June 23, 2026 unveiling at the Volcano Engine FORCE conference; single continuous 30s clips without stitching.
  • China's AI Video Battle Picks Sides: Seedance 2.5 Stays Closed, MiniMax H3 Goes Open — Tech Times — competitive context for the Seedance 2.5 vs MiniMax H3 comparison.
  • ByteDance reportedly pauses global launch of its Seedance 2.0 video generator — TechCrunch — background on the 2.0 generation this release supersedes.
  • Seedance 2.0 audio input guide — Volcano Engine — official guidance on choosing clean reference audio, the lineage the 2.5 audio-only mode builds on.
All Posts

Author

avatar for Vogoo AI Team
Vogoo AI Team

Categories

  • Guide
What is audio-only reference?Audio reference specs: what you can uploadFive official cases, from pure audio to full multimodalCase 1 — Audio only: two geckos plan a jailbreakCase 2 — Rhythm transfer: the headphone adCase 3 — Three modalities at once: the walking one-shotCase 4 — Music-driven multi-character: the ballroomCase 5 — Music over a 3D white model: the spaceshipBuilding your own audio + image / video promptsWhat makes this differentWhere audio-driven generation fitsFAQWhat audio formats does Seedance 2.5 accept as reference?Can Seedance 2.5 generate a video from only a song or audio clip?Does audio reference handle lip sync?Do audio references reduce how much video I can reference?How is audio to video AI different from adding music in an editor?Where can I try Seedance 2.5's audio features?The bottom lineSources

More Posts

Twitter Header Size: 1500×500 & the Safe Area Explained
Guide

Twitter Header Size: 1500×500 & the Safe Area Explained

What is the right Twitter header size? The official 1500×500 spec, the avatar safe area you must avoid, and how to export a banner that stays sharp.

avatar for Vogoo AI Team
Vogoo AI Team
2026/08/13
Seedance 2.5 vs Seedance 2.0: Every Upgrade Compared
News

Seedance 2.5 vs Seedance 2.0: Every Upgrade Compared

Seedance 2.5 vs Seedance 2.0 compared: 30s one-shot generation, 50 references, audio-only input, precise video editing — and who should actually switch.

avatar for Vogoo AI Team
Vogoo AI Team
2026/08/06
Text to Image vs Image to Image: Which Should You Use?
Product

Text to Image vs Image to Image: Which Should You Use?

Text to image vs image to image, explained: how each one works, what they're best at, a decision table for picking one, and when to combine both.

avatar for Vogoo AI Team
Vogoo AI Team
2026/07/28

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates

Vogoo.ai

The free AI picture generator that exports at exact pixel image sizes — start from a prompt template, post without cropping.

X (Twitter)
Product
  • Text to Video
  • Image to Video
  • Text to Image
  • Image to Image
  • Text to Music
  • AI Album Cover Generator
  • AI Flyer Generator
  • AI Wallpaper Generator
  • AI Christmas Photo Generator
  • Pricing
AI Models
  • Reve 2.0
  • Reve 2.1
  • Nano Banana 2 Lite
  • Seedream 5.0 Pro
  • Qwen Image 3.0
  • Seedream 5.0 Lite
  • Gemini Omni Flash
  • Flux 3
  • LTX 2.5
  • MiniMax H3
  • HappyHorse 1.0
  • Seedance 2.5
  • Seedance 2.0 Mini
  • Wan 2.7
  • iLoveSong AI
Company
  • About
  • Contact us
  • Social Media Visuals Guide
Legal
  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Vogoo.ai. All Rights Reserved.Vogoo AI Guide on Gamma