Home/RESEARCH GUIDE
RESEARCH GUIDE

A guide to ByteDance’s published Seed audio research

A source-led introduction to Seed-TTS for speech generation and Seed-Music for controlled music generation.

Source-led editorial guideUpdated July 28, 2026
THE SHORT ANSWER

ByteDance’s published audio research includes Seed-TTS, a family of versatile speech-generation models, and Seed-Music, a framework for controlled music generation. Together, the papers document advances in voice generation, speech editing and multimodal control of vocal music.

SEED-TTS

Published speech-generation capabilities

The Seed-TTS technical report describes zero-shot speech in-context learning, also called zero-shot voice continuation, together with speech editing and controlled speech generation [1].

The report presents objective evaluations and human evaluations across several speech-generation tasks [1].

SEED-MUSIC

Published music-generation capabilities

The Seed-Music paper presents controlled vocal-music generation from multimodal inputs [2].

Its documented controls include style descriptions, audio references, musical scores and voice prompts [2].

EDITORIAL METHOD

How this site turns sources into useful guidance

Use the exact model name

Each capability stays connected to the model named in its primary source.

Link the primary document

Technical claims point to an official paper, model card, API reference or product guide.

Name the implementation

Provider features are attributed to the provider and release that documents them.

Keep a test record

A complete comparison record should include prompts, settings, output samples and evaluation criteria.

FAQ

Questions people ask before choosing

What is Seed-TTS?

Seed-TTS is ByteDance’s published family of speech-generation models, covering speech in-context learning, speech editing and controlled generation.

What is Seed-Music?

Seed-Music is ByteDance’s published framework for high-quality, controlled music generation from multimodal inputs.

Which inputs appear in the Seed-Music paper?

The paper describes style descriptions, audio references, musical scores and voice prompts.

How are commercial terms checked?

Use the license and service terms attached to the exact model and provider selected for production.

SOURCES

Evidence used for this guide

Product facts are drawn from the official material below. Editorial methods and recommendations are labeled throughout the guide.

  1. Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsByteDance technical report supporting the speech-generation statements marked [1].
  2. Seed-Music: A Unified Framework for High Quality and Controlled Music GenerationByteDance research paper supporting the music-generation statements marked [2].