Recently, ByteDance officially launched the audio creation model Seed Audio 1.0, marking a new stage in AI audio generation technology, moving from a single speech synthesis phase to a complete sound scene creation. The model is now available at the Volcano Fangzhou Experience Center and is open for testing by creators.
For a long time, the production of film and television-level audio content relied on multiple models separately generating vocals, sound effects, and ambient sounds, then manually splicing and mixing them, which was a lengthy process and difficult to ensure overall narrative consistency. The core breakthrough of Seed Audio 1.0 lies in that it does not simply concatenate materials but jointly models various audio elements under a unified framework, generating "complete sound works" that serve storytelling end-to-end.

According to the official introduction, Seed Audio 1.0 has three core capabilities. First, precise temporal-spatial arrangement, supporting control of dialogue and sound effects entry times with 100-millisecond accuracy along the timeline, perfectly suitable for video dubbing and advertising production. Second, stable voice performance, supporting zero-shot generation and long audio extension, maintaining character voice consistency while naturally expressing different emotions such as anger and joy, even allowing the same voice to perform multiple characters. Third, fluent multilingual support, covering more than 20 languages including Chinese, English, and Japanese, ensuring that the voice conforms to local expression rhythms and stress habits in different languages.
Evaluation data shows that the model's audio availability in nine common creative scenarios has exceeded 90%, and the naturalness MOS score for multilingual generation is generally above 4 points (excellent level).

From "being able to speak" to "being able to create," the release of Seed Audio 1.0 reduces the threshold for producing high-quality audio content. In the future, the team plans to further integrate multimodal inputs such as video references and explore controllable translation technology, continuously optimizing long audio and track generation capabilities, helping creators efficiently transform their mental sound concepts into audible works.
Project Homepage:
https://seed.bytedance.com/seedaudio1_0
Experience Access:
Volcano Fangzhou Experience Center - Login - Select Voice Model - Voice Synthesis - Doubao - Audio Generation - 1.0
