Skip to main content
POST
v1/audio/speech (TTS + Music)

Overview

Synthesize speech (TTS) or generate music with a single request. The OpenAI-compatible POST /v1/audio/speech endpoint returns audio bytes directly in the HTTP response. Optional chunked streaming is available with stream: true. In streaming mode, audio bytes are delivered progressively as generated, which reduces time-to-first-byte (TTFB) for real-time playback. Default behavior is unchanged: omit stream (or set false) to receive one buffered audio file after generation completes. When you use a music model (for example Minimax-Music-02), the input field is treated as a music prompt (not text to speak) and voice is ignored. For a dedicated music guide and model list, see api-reference/music-generation.mdx.

Endpoint

  • Method/Path: POST https://nano-gpt.com/api/v1/audio/speech
  • Auth: Authorization: Bearer <API_KEY>
  • Required header: Content-Type: application/json
  • You may see older examples using POST https://nano-gpt.com/api/v1/speech. Prefer /api/v1/audio/speech for OpenAI SDK compatibility.

Request Parameters

Notes:
  • Some provider-backed models may support additional fields; accepted parameters vary by model.
  • For unsupported models, stream: true is ignored and the endpoint returns the normal buffered response.

Streaming Support (TTS Models)

All other TTS models (Gemini, Inworld, Kokoro, Qwen, MiniMax, and others) ignore stream and return buffered responses.

Response Behavior

Non-streaming (default)

  • Status: 200 OK
  • Body: complete audio file returned once generation finishes
  • Headers: typically includes Content-Length

Streaming (stream: true)

  • Status: 200 OK
  • Body: audio bytes arrive progressively in chunks
  • Headers: no Content-Length; uses Transfer-Encoding: chunked
  • Client can start playback/processing as soon as the first chunk arrives
  • Content-Type by provider:
    • OpenAI TTS models: matches the selected output format (for example audio/mpeg, audio/wav, audio/opus, audio/flac, audio/aac, audio/pcm)
    • ElevenLabs TTS models: always audio/mpeg

Errors

  • If an error happens before streaming starts, the API returns the standard OpenAI-style JSON error envelope:
  • If streaming fails after bytes have started, the connection is terminated and clients may receive partial/corrupt audio.
Common error types: invalid_model, invalid_voice, unsupported_format, input_too_long, rate_limit_exceeded.

Examples

Non-streaming request (unchanged)

Streaming request (cURL)

OpenAI Node.js SDK (streaming)

OpenAI Python SDK (streaming)

Compare TTFB (streaming vs non-streaming)

Notes & Limits

  • Max input length: depends on model; measured in characters or tokens. For short, interactive prompts, prefer under ~1-2k characters.
  • Typical latency: scales with input length and output format; compressed formats like mp3 are often faster than wav.
  • Usage metering: billed by input characters for TTS models; output file size does not affect billing.

Audio Format Support by Provider

Voices

  • Voice IDs vary by model/provider. See model-specific voices on Text-to-Speech: api-reference/text-to-speech.mdx.
  • If a voices listing endpoint is available (for example GET /v1/voices), it returns available voice IDs and metadata (language coverage, gender/pitch, sample links).

Errors & Troubleshooting

  • invalid_model, invalid_voice, unsupported_format: Verify model, voice, and response_format.
  • input_too_long: Reduce length; split long text into chunks and stitch audio client-side.
  • rate_limit_exceeded: Exponential backoff; retry after the window resets.
  • Network/client tips: set Accept to your preferred audio type and write raw response bytes directly to a file/stream.

Security

  • Do not expose API keys in browsers. Proxy via your server.
  • Redact PII in logs; avoid logging raw text/audio in production.
  • Rate-limit public routes.

Pricing, Quotas, and Rate Limits

  • Billing is based on input character count, not output audio size.
  • For streaming requests, billing is recorded after the first audio chunk is confirmed. If the upstream provider fails before any audio is produced, no charge is applied.
  • If the client disconnects mid-stream after audio starts, the charge still applies because generation already began upstream.
  • Rate limits: per-minute/day caps; contact support to request increases. See api-reference/miscellaneous/pricing.mdx and api-reference/miscellaneous/rate-limits.mdx.

Migration from Job-based TTS

Already using the async POST /tts + GET /tts/status flow?
  • When to switch: choose v1/audio/speech for short prompts, low latency, and direct playback; keep job-based TTS for long/batch generation and webhook workflows.
  • Parameter mapping: text -> input, voice stays voice, and output format can be requested with response_format when supported.
  • Retries/timeouts: v1/audio/speech returns inline; implement client-side timeouts and simple retries on 5xx.

See Also

  • Async/job-based TTS: api-reference/endpoint/tts.mdx
  • TTS Status polling: api-reference/endpoint/tts-status.mdx