What the Seed Audio 1.0 label means on this site
This site uses “Seed Audio 1.0” as an editorial label for navigating ByteDance's early published Seed-TTS and Seed-Music work. The cited papers name the systems Seed-TTS and Seed-Music; they do not define a unified model called “Seed Audio 1.0” [1][2].
For that reason, the capabilities below are attributed to the exact system documented by each source. They should not be read as evidence that every capability belongs to one combined model.
Documented speech-generation capabilities
The Seed-TTS technical report introduces a family of large-scale autoregressive text-to-speech models. It reports speech in-context learning and evaluates speaker similarity and naturalness with objective and subjective measures [1].
The report also describes control over speech attributes such as emotion, expressive and diverse speech generation, and a diffusion-based non-autoregressive variant called Seed-TTS_DiT. The authors demonstrate the latter's use for speech editing [1].
Speech in-context learning
The report evaluates continuation from speech context, including speaker similarity and naturalness [1].
Attribute control
The paper explicitly discusses control over speech attributes such as emotion [1].
Expressive speech
Seed-TTS is described as generating expressive and diverse speech for speakers in the wild [1].
Speech editing
The report demonstrates speech editing with the diffusion-based Seed-TTS_DiT variant [1].
Documented music-generation and editing capabilities
The Seed-Music paper introduces a suite of music-generation systems that combines autoregressive language modeling and diffusion approaches. It identifies two principal workflows: controlled music generation and post-production editing [2].
For controlled vocal-music generation, the paper lists multimodal inputs including style descriptions, audio references, musical scores, and voice prompts. For post-production, it describes interactive editing of lyrics and vocal melodies in generated audio [2].
ByteDance's official Seed-Music page additionally describes expressive vocals in multiple languages, fine-grained note-level editing, and zero-shot singing voice conversion using a short singing or speech recording [3].
| Documented area | Capability | Source |
|---|---|---|
| Generation | Controlled vocal-music generation | Seed-Music paper [2] |
| Inputs | Style descriptions, audio references, musical scores, and voice prompts | Seed-Music paper [2] |
| Editing | Editing lyrics and vocal melodies in generated audio | Seed-Music paper [2] |
| Voice conversion | Zero-shot singing voice conversion from a short singing or speech recording | Official Seed-Music page [3] |
What these sources do not establish
The cited sources do not establish “Seed Audio 1.0” as the official name of a single product, nor do they describe an official “Seed Audio 1.0 versus Seed Audio 2.0” version relationship.
They also do not show that Seed-TTS and Seed-Music run as one combined model. Any implementation, availability, pricing, license, or commercial-use claim must be checked against the terms and documentation for the exact service through which a model is accessed.
Questions people ask before choosing
Is Seed Audio 1.0 an official ByteDance model name?
The sources cited here do not define a single model by that name. This site uses it only as an editorial navigation label for the early Seed-TTS and Seed-Music research covered on this page.
What is Seed-TTS?
Seed-TTS is a family of speech-generation models documented in ByteDance's technical report, including autoregressive models and a diffusion-based Seed-TTS_DiT variant [1].
What is Seed-Music?
Seed-Music is a suite of music-generation systems for controlled generation and post-production editing, documented in its paper and official project page [2][3].
Are Seed-TTS and Seed-Music one combined model?
The cited sources document them separately and do not establish that they are one combined model.
Can these models be used commercially?
This page does not make a commercial-use claim. Check the license and service terms attached to the exact model and provider you intend to use.
Evidence used for this guide
Product facts are drawn from the official material below. Editorial methods and recommendations are labeled throughout the guide.
- [1] Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsPrimary technical report supporting the Seed-TTS descriptions on this page.
- [2] Seed-Music: A Unified Framework for High Quality and Controlled Music GenerationPrimary paper supporting the Seed-Music generation, input, and editing descriptions.
- [3] Hi, Seed-Music — ByteDance SeedOfficial project page supporting the multilingual vocals, note-level editing, and zero-shot singing voice-conversion descriptions.