Home/SOURCE-LED GUIDE
SOURCE-LED GUIDE

Seed Audio 1.0: a guide to the early Seed-TTS and Seed-Music research

A carefully sourced overview of ByteDance's published Seed-TTS speech-generation and Seed-Music controlled-music research.

Source-led editorial guideUpdated August 3, 2026
THE SHORT ANSWER

“Seed Audio 1.0” is an editorial navigation label used by this site; it is not presented as the official name of a single ByteDance model or product. The documented subjects on this page are Seed-TTS, a family of speech-generation models, and Seed-Music, a suite of controlled music-generation systems [1][2][3].

NAMING

What the Seed Audio 1.0 label means on this site

This site uses “Seed Audio 1.0” as an editorial label for navigating ByteDance's early published Seed-TTS and Seed-Music work. The cited papers name the systems Seed-TTS and Seed-Music; they do not define a unified model called “Seed Audio 1.0” [1][2].

For that reason, the capabilities below are attributed to the exact system documented by each source. They should not be read as evidence that every capability belongs to one combined model.

SEED-TTS

Documented speech-generation capabilities

The Seed-TTS technical report introduces a family of large-scale autoregressive text-to-speech models. It reports speech in-context learning and evaluates speaker similarity and naturalness with objective and subjective measures [1].

The report also describes control over speech attributes such as emotion, expressive and diverse speech generation, and a diffusion-based non-autoregressive variant called Seed-TTS_DiT. The authors demonstrate the latter's use for speech editing [1].

Speech in-context learning

The report evaluates continuation from speech context, including speaker similarity and naturalness [1].

Attribute control

The paper explicitly discusses control over speech attributes such as emotion [1].

Expressive speech

Seed-TTS is described as generating expressive and diverse speech for speakers in the wild [1].

Speech editing

The report demonstrates speech editing with the diffusion-based Seed-TTS_DiT variant [1].

SEED-MUSIC

Documented music-generation and editing capabilities

The Seed-Music paper introduces a suite of music-generation systems that combines autoregressive language modeling and diffusion approaches. It identifies two principal workflows: controlled music generation and post-production editing [2].

For controlled vocal-music generation, the paper lists multimodal inputs including style descriptions, audio references, musical scores, and voice prompts. For post-production, it describes interactive editing of lyrics and vocal melodies in generated audio [2].

ByteDance's official Seed-Music page additionally describes expressive vocals in multiple languages, fine-grained note-level editing, and zero-shot singing voice conversion using a short singing or speech recording [3].

Documented areaCapabilitySource
GenerationControlled vocal-music generationSeed-Music paper [2]
InputsStyle descriptions, audio references, musical scores, and voice promptsSeed-Music paper [2]
EditingEditing lyrics and vocal melodies in generated audioSeed-Music paper [2]
Voice conversionZero-shot singing voice conversion from a short singing or speech recordingOfficial Seed-Music page [3]
SCOPE

What these sources do not establish

The cited sources do not establish “Seed Audio 1.0” as the official name of a single product, nor do they describe an official “Seed Audio 1.0 versus Seed Audio 2.0” version relationship.

They also do not show that Seed-TTS and Seed-Music run as one combined model. Any implementation, availability, pricing, license, or commercial-use claim must be checked against the terms and documentation for the exact service through which a model is accessed.

FAQ

Questions people ask before choosing

Is Seed Audio 1.0 an official ByteDance model name?

The sources cited here do not define a single model by that name. This site uses it only as an editorial navigation label for the early Seed-TTS and Seed-Music research covered on this page.

What is Seed-TTS?

Seed-TTS is a family of speech-generation models documented in ByteDance's technical report, including autoregressive models and a diffusion-based Seed-TTS_DiT variant [1].

What is Seed-Music?

Seed-Music is a suite of music-generation systems for controlled generation and post-production editing, documented in its paper and official project page [2][3].

Are Seed-TTS and Seed-Music one combined model?

The cited sources document them separately and do not establish that they are one combined model.

Can these models be used commercially?

This page does not make a commercial-use claim. Check the license and service terms attached to the exact model and provider you intend to use.

SOURCES

Evidence used for this guide

Product facts are drawn from the official material below. Editorial methods and recommendations are labeled throughout the guide.

  1. [1] Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsPrimary technical report supporting the Seed-TTS descriptions on this page.
  2. [2] Seed-Music: A Unified Framework for High Quality and Controlled Music GenerationPrimary paper supporting the Seed-Music generation, input, and editing descriptions.
  3. [3] Hi, Seed-Music — ByteDance SeedOfficial project page supporting the multilingual vocals, note-level editing, and zero-shot singing voice-conversion descriptions.