SaydiVoice API reference
This page renders interactively with JavaScript. The complete reference is below, and is also available as plain text at /llms-full.txt.
Base URL https://voice.saydi.ai/api. Authenticate with Authorization: Bearer sv_live_….
Introduction
Text to speech and voice cloning over one REST API.
SaydiVoice turns text into speech and clones a voice from a short reference clip. One bearer token, audio bytes out.
What you can call — Text to speech (POST /v1/audio/speech, OpenAI-compatible), voice cloning (POST /tts/clone), transcription (POST /voice/transcribe), plus the /samples catalogue.
Quick Start
Create a key and get your first audio file.
- Create an API key — Open /developers/#/keys, create a key and copy the secret (shown once).
- Store it safely — Put it in an environment variable. Never commit it or ship it in client-side code.
- Make your first request — POST to /v1/audio/speech and write the response body to a file.
cURL
export SAYDI_API_KEY="sv_live_REPLACE_WITH_YOUR_KEY"
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"vi-adam","input":"Xin chào, đây là API của SaydiVoice."}' \
--output hello.mp3
# Inspect what you were charged (characters) and how long it took:
# -D - prints the response headers, e.g. X-OV-Chars: 41
Python
import os, requests
API_KEY = os.environ["SAYDI_API_KEY"]
BASE = "https://voice.saydi.ai/api"
r = requests.post(
f"{BASE}/v1/audio/speech",
headers={"Authorization": f"Bearer {API_KEY}"},
json={
"model": "tts-1",
"voice": "vi-adam",
"input": "Xin chào, đây là API của SaydiVoice.",
},
timeout=120,
)
r.raise_for_status()
open("hello.mp3", "wb").write(r.content)
print("chars charged:", r.headers.get("X-OV-Chars"))
Node.js
import { writeFile } from "node:fs/promises";
const API_KEY = process.env.SAYDI_API_KEY;
const BASE = "https://voice.saydi.ai/api";
const res = await fetch(`${BASE}/v1/audio/speech`, {
method: "POST",
headers: {
Authorization: `Bearer ${API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "tts-1",
voice: "vi-adam",
input: "Xin chào, đây là API của SaydiVoice.",
}),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
await writeFile("hello.mp3", Buffer.from(await res.arrayBuffer()));
console.log("chars charged:", res.headers.get("X-OV-Chars"));
Expect audio bytes — Success returns audio bytes (audio/mpeg or audio/wav). A JSON response means an error.
Authentication
Authenticate with one API key.
Send the key as Authorization: Bearer sv_live_.... Never put it in URLs, logs or client-side code.
Request headers
Authorization: Bearer sv_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Content-Type: application/json
Manage keys
Create, list, rotate and revoke API keys.
These routes manage keys and require a signed-in user token, not an API key.
GET /api-keys — List keys
Returns every key on the account. Secrets and hashes are never included.
cURL
curl -sS "https://voice.saydi.ai/api/api-keys" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — The account's keys, newest first.
{
"ok": true,
"max_keys": 20,
"keys": [
{
"id": "key_1f0c9a2b7d3e4f56",
"name": "Production",
"prefix": "sv_live_Ab12Cd",
"scopes": ["tts"],
"status": "active",
"created_at": 1758844800.0,
"last_used_at": 1758848400.0,
"expires_at": null,
"revoked_at": null
}
]
}
POST /api-keys — Create a key
Creates a key and returns the secret exactly once, in the key field.
| Name | Type | Required | Description |
|---|
name | string | no | A label to recognise the key later. Defaults to "Default key". |
scopes | string[] | no | Reserved for later releases. Defaults to ["tts"]. |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/api-keys" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"Production"}'
201 — Created. Store key now — it cannot be read again.
{
"ok": true,
"id": "key_1f0c9a2b7d3e4f56",
"name": "Production",
"prefix": "sv_live_Ab12Cd",
"scopes": ["tts"],
"status": "active",
"created_at": 1758844800.0,
"key": "sv_live_Ab12CdEf34Gh56Ij78Kl90Mn12Op34Qr"
}
409 — Per-account key limit reached. Revoke an unused key first.
{"detail": "API key limit reached (20). Revoke an unused key first."}
DELETE /api-keys/{key_id} — Revoke a key
Revokes immediately. The key stops authenticating on the next request.
cURL
curl -sS -X DELETE "https://voice.saydi.ai/api/api-keys/key_1f0c9a2b7d3e4f56" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Revoked.
{"ok": true, "id": "key_1f0c9a2b7d3e4f56", "status": "revoked"}
404 — No such key on this account.
{"detail": "API key not found."}
POST /api-keys/{key_id}/rotate — Rotate a key
Issues a new secret for the same key id. The previous secret stops working immediately.
Keeps the key id and name, replaces only the secret.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/api-keys/KEY_ID/rotate" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json"
GET /api-keys/status — Key management availability
Public readiness probe: whether key management is deployed, the per-account cap, and the known scopes.
cURL
curl -sS "https://voice.saydi.ai/api/api-keys/status"
200 — Readiness snapshot.
{"available": true, "max_keys": 20, "scopes": ["tts", "voices"]}
POST /v1/audio/speech
OpenAI-compatible text to speech.
OpenAI-compatible endpoint: same path, same body fields, same audio-bytes response, same error envelope. Read the table below before swapping base_url — two fields behave differently on purpose.
| Part | OpenAI | SaydiVoice |
|---|
| Path, auth, JSON body, audio bytes | POST /v1/audio/speech | Identical |
| model | tts-1, tts-1-hd, gpt-4o-mini-tts | tts-1 → 16 steps, tts-1-hd → 32 steps; any other id is accepted and uses the server default |
| voice | 13 built-in voices, or {id} for a custom voice | A SaydiVoice voice. The 13 OpenAI names are accepted but ALL resolve to one default (see the callout below); the {id} object form is rejected |
| instructions | Free-form style sentence on a fixed voice (gpt-4o-mini-tts) | Voice-design mode with a PRESET vocabulary (gender, age, pitch, accent, whisper). A free-form sentence is 400 with error.param "instructions"; so is combining it with a SaydiVoice voice |
| response_format | mp3, opus, aac, flac, wav, pcm | Same six; aac is served as mp3 |
| stream_format | audio (chunked), sse (audio events) | audio is accepted but the file arrives in one response; sse is rejected with 400 |
| Errors | {"error": {...}}, 400 for a bad field | Same envelope here; a bad field keeps status 422. Native endpoints keep {"detail": ...} |
| Name | Type | Required | Description |
|---|
input | string | yes | Text to speak. Schema maximum 4096 characters; your plan's per-request character cap also applies (Free: 3000), and exceeding it is a 429 with error.code "insufficient_quota". Use POST /tts for longer text. |
voice | string | yes | A voice_id from GET /samples (legacy name still accepted), a registered clone voice from POST /v1/voices (vc_...), or an OpenAI built-in name (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar). An OpenAI name does NOT select that voice: none exist here, so they all resolve to the alias default (vi-adam unless OMNIVOICE_OPENAI_VOICE_DEFAULT says otherwise) and the response shows X-OV-Voice-Source: openai-alias-default. Any other unknown name falls back to auto-voice. X-OV-Voice always reports the voice that actually spoke, or empty for auto. |
model | string | no | "tts-1" = 16 diffusion steps (faster), "tts-1-hd" = 32 steps (higher quality). Omitted = server default. |
response_format | string | no | mp3 (default), opus, aac, flac, wav, pcm. aac is currently served as mp3. |
speed | number | no | Speech speed, 0.25–4.0. Default 1.0. |
instructions | string | no | Style instruction, mapped to this server's voice-design prompt (same engine as instruct on POST /tts). The vocabulary is PRESET — gender, age, pitch, accent, whisper — for example "female, british accent, whisper". A free-form style sentence (what OpenAI's gpt-4o-mini-tts accepts, e.g. "Speak in a cheerful and positive tone.") is refused with 400, error.param "instructions", and the message lists the valid items. It also cannot be combined with a SaydiVoice voice: design mode builds the speaker from the instruction and has no reference clip. |
stream_format | string | no | Compatibility field. 'audio' is accepted and the whole file arrives in one response; 'sse' is rejected with 400 because audio events are not implemented. |
OpenAI voice names are accepted, but they do not select a voice — None of alloy/ash/ballad/coral/echo/fable/nova/onyx/sage/shimmer/verse/marin/cedar exist in this catalogue, so every one of them resolves to a single default voice (vi-adam unless OMNIVOICE_OPENAI_VOICE_DEFAULT says otherwise). The response says so: X-OV-Voice-Source: openai-alias-default plus X-OV-Voice-Hint. To choose a voice, pass a voice_id from GET /samples — X-OV-Voice then reports the voice that actually spoke, and saydivoice as the source.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1-hd",
"voice": "vi-adam",
"input": "Higher quality, 32 diffusion steps.",
"response_format": "wav",
"speed": 1.0
}' \
--output speech.wav
OpenAI SDK (base_url swap)
from openai import OpenAI
client = OpenAI(
api_key=os.environ["SAYDI_API_KEY"],
base_url="https://voice.saydi.ai/api/v1", # only this line changes
)
# The body is unchanged. `voice` is a SaydiVoice voice_id from GET /samples:
# OpenAI's built-in names are accepted but all resolve to one default voice.
with client.audio.speech.with_streaming_response.create(
model="tts-1",
voice="vi-adam",
input="Drop-in replacement.",
) as response:
response.stream_to_file("speech.mp3")
Metering headers — Each audio response reports cost and timing: X-OV-Chars, X-Audio-Duration, X-Generation-Time, X-RTF.
| Header | Meaning |
|---|
| X-OV-Chars | Characters charged for this request |
| X-OV-Voice | Voice actually used (after fallback) |
| X-Audio-Duration | Generated audio length in seconds |
| X-Generation-Time | Compute time in seconds |
| X-RTF | Real-time factor: generation time / audio duration |
POST /tts
Full-control text to speech.
Use this endpoint for long text, fixed duration, auto pauses, voice design or sampling controls.
Voice selection — three modes
| Mode | Set this | Result |
|---|
| Voice cloning | sample (+ optional ref_text) | Speaks in a catalogue voice copied from its reference clip |
| Voice design | instruct | The model creates a voice matching a text description |
| Auto voice | neither | The model picks a voice for the text |
| Name | Type | Required | Description |
|---|
text | string | yes | Text to synthesize. |
sample | string | no | Voice_id from GET /samples (legacy name still accepted). Enables cloning mode. |
voice_id | string | no | A registered clone voice from POST /v1/voices (e.g. vc_9f2c...). No reference audio needed; the server uses the stored clip and transcript. |
instruct | string | no | Voice design description, e.g. "female, low pitch, british accent". Ignored when sample is set. |
ref_text | string | no | Override the reference transcript for the sample. Improves cloning accuracy. |
lang | string | no | Language hint (ISO 639-1). Used for text normalization and pronunciation. |
breaks | object | no | Auto pauses in ms per punctuation type: {"sentence":450,"comma":250,"semicolon":300,"paragraph":600}. |
duration | number | no | Fixed output duration in seconds (max 600). Overrides speed. Useful to fit a time slot. |
speed | number | no | Speed factor, 0–10. Ignored when duration is set. |
num_step | integer | no | Diffusion steps 1–256. Default 32; 16 is faster. |
guidance_scale | number | no | Classifier-free guidance, 0–50. Default 2.0. |
t_shift | number | no | Noise-schedule shift, 0–1. Default 0.1. |
denoise | boolean | no | Prepend the denoising token. Default true. |
preprocess_prompt | boolean | no | Clean the reference clip before cloning. Default true. |
postprocess_output | boolean | no | Trim long silences from the output. Default true. |
audio_chunk_duration | number | no | Target chunk length in seconds for long text (max 120). Default 15. |
audio_chunk_threshold | number | no | Estimated duration that triggers chunking (max 600). Default 30. |
output_format | string | no | wav, mp3, flac or ogg. Default from server config. |
Long text with pauses and a fixed duration
curl -sS -X POST "https://voice.saydi.ai/api/tts" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Câu thứ nhất. Câu thứ hai, có dấu phẩy.",
"sample": "vi-adam",
"lang": "vi",
"breaks": {"sentence": 450, "comma": 250},
"num_step": 32,
"output_format": "mp3"
}' \
--output long.mp3
Voice design (no reference clip)
curl -sS -X POST "https://voice.saydi.ai/api/tts" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "A calm narrator for a documentary.",
"instruct": "male, middle-aged, low pitch, american accent",
"lang": "en"
}' \
--output designed.mp3
POST /tts/file
Generate speech from a .txt file.
Send the script as a UTF-8 .txt file instead of embedding a long string in JSON.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/tts/file" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "text_file=@script.txt" \
-F "sample=vi-adam" \
-F "lang=vi" \
-F "output_format=mp3" \
--output script.mp3
Accepts every field of POST /tts; use text_file instead of text.
GET /samples
List voices and metadata.
Pass voice_id as voice (OpenAI-compatible) or sample (native). Legacy name still works; use display_name for UI labels.
| Name | Type | Required | Description |
|---|
lang | string | no | Return only voices whose primary language matches, e.g. "vi". |
limit | integer | no | Return at most this many voices, after sorting by usage. |
cURL
curl -sS "https://voice.saydi.ai/api/samples?lang=vi&limit=5" \
-H "Authorization: Bearer $SAYDI_API_KEY"
200 — An array of voice objects.
[
{
"name": "adam-11labs-vi",
"voice_id": "vi-adam",
"audio_file": "adam-11labs-vi.wav",
"has_transcript": true,
"transcript_preview": "Xin chào, tôi là Adam...",
"display_name": "Adam — Giọng hot tiktok",
"gender": "male",
"age": "young",
"accent": "vietnamese",
"locale": "vi-VN",
"language": "vi",
"languages": ["vi"],
"use_case": "conversational",
"category": "high_quality",
"featured": true,
"usage_count": 12840
}
]
GET /samples/{name}/audio — Preview a voice
Streams the reference clip for a voice, so you can audition it before generating.
cURL
curl -sS "https://voice.saydi.ai/api/samples/vi-adam/audio" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
--output preview.wav
GET /v1/models — List models
OpenAI-compatible model list: tts-1 (16 steps) and tts-1-hd (32 steps).
cURL
curl -sS "https://voice.saydi.ai/api/v1/models" \
-H "Authorization: Bearer $SAYDI_API_KEY"
POST /v1/voices
Register a voice once, reuse it by voice_id.
Upload a clip and a name. SaydiVoice checks the audio, picks its best continuous stretch of speech, transcribes it, and removes background music when it finds any — then returns a stable voice_id. Use that id in POST /v1/audio/speech or POST /tts: no clip to re-upload, and the reference is cached server-side so later generations start faster. Voices registered here also appear in Studio → My Voices, and voices you create in Studio appear here.
POST /v1/voices — Register a voice
Multipart upload. Returns the voice object; status is ready when it is usable right away.
| Name | Type | Required | Description |
|---|
file | file | yes | Reference clip: WAV, MP3, M4A, OGG, FLAC or WebM. Any length up to 5 minutes; the best stretch is selected for you. |
name | string | yes | Display name. Not unique — identify the voice by id, never by name. |
consent | string | yes | Confirm you have the right to clone this voice. Send true, or JSON {"accepted": true}. Recorded with your IP and the terms version. |
language | string | no | ISO-639-1 hint (e.g. vi). Also picks the language of the preview sentence. |
ref_text | string | no | Transcript of the clip. Omit it and we transcribe for you; send it (when your clip is already prepared) to skip that step. |
separate | string | no | auto (default) separates background music only when detected; true always does; false never does. |
Idempotency-Key | string | no | Header. Retrying with the same value returns the voice that already exists instead of creating a second one — so a timeout never costs you a clone slot. |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/v1/voices" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "file=@my-recording.m4a" \
-F "name=Giọng của tôi" \
-F "language=vi" \
-F "consent=true"
# -> {"id":"vc_9f2c...","status":"ready","ref_text":"...","quality":{...}}
# Then speak with it — no clip to upload:
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"vc_9f2c...","input":"Xin chào từ giọng của tôi."}' \
--output hello.mp3
Python
import os, requests
key = os.environ["SAYDI_API_KEY"]
base = "https://voice.saydi.ai/api"
with open("my-recording.m4a", "rb") as fh:
voice = requests.post(
f"{base}/v1/voices",
headers={"Authorization": f"Bearer {key}"},
files={"file": ("my-recording.m4a", fh, "audio/mp4")},
data={"name": "Giọng của tôi", "language": "vi", "consent": "true"},
timeout=300,
)
voice.raise_for_status()
voice = voice.json()
print(voice["id"], voice["status"], voice.get("quality", {}).get("warnings"))
audio = requests.post(
f"{base}/v1/audio/speech",
headers={"Authorization": f"Bearer {key}"},
json={"model": "tts-1", "voice": voice["id"], "input": "Xin chào từ giọng của tôi."},
timeout=180,
)
audio.raise_for_status()
open("hello.mp3", "wb").write(audio.content)
201 — The voice was created.
{
"id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12",
"object": "voice",
"name": "Giọng của tôi",
"status": "ready",
"language": "vi",
"ref_seconds": 7.4,
"ref_text": "Xin chào, đây là đoạn nói trong clip mẫu.",
"quality": {
"speech_seconds": 6.9,
"snr_db": 26.4,
"trimmed_seconds": 42.1,
"music_detected": false,
"separated": false,
"warnings": []
},
"consent": {"terms_version": "voice-terms-2026-01"},
"preview_url": "/v1/voices/vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12/preview"
}
200 — An idempotent retry: this voice already existed and is returned unchanged.
{
"id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12",
"status": "ready",
"idempotent_replay": true
}
processing vs ready — ready means usable now — that is the normal case for a short clean clip. processing means the clip had music under the voice, so separation is still running: poll GET /v1/voices/{voice_id} (usually seconds) before synthesising. A failed voice frees its slot and carries an error.code you can act on. Two errors are worth handling separately: 503 storage_unavailable is temporary — retry with backoff; 410 ref_missing means the stored audio is gone for good, so register the voice again (it does not use an extra slot).
GET /v1/voices — List and manage voices
List your voices, fetch one, rename it, delete it, or hear a preview.
| Endpoint | What it does |
|---|
GET /v1/voices | This account's voices, newest first, with usage (slots used/limit). Filter with ?status=ready. |
GET /v1/voices/{voice_id} | One voice: status, ref_text and the quality report. |
PATCH /v1/voices/{voice_id} | Rename: JSON {"name": "..."}. The id never changes. |
DELETE /v1/voices/{voice_id} | Delete the voice and its stored audio, freeing its clone slot. Immediate. |
GET /v1/voices/{voice_id}/preview | MP3 of a short sentence in that voice. Rendered once and cached. Not charged against your characters. |
cURL
curl -sS "https://voice.saydi.ai/api/v1/voices" \
-H "Authorization: Bearer $SAYDI_API_KEY"
curl -sS -X PATCH "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Giọng nam trầm"}'
curl -sS "https://voice.saydi.ai/api/v1/voices/$VOICE_ID/preview" \
-H "Authorization: Bearer $SAYDI_API_KEY" --output preview.mp3
curl -sS -X DELETE "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \
-H "Authorization: Bearer $SAYDI_API_KEY"
Consent and slot limits — Only register voices you have the right to use; the API refuses a registration without consent. Each plan includes a number of clone slots (clone_slots) — creating a voice beyond it returns 402 with code: voice_limit_reached. A failed registration does not consume a slot.
POST /tts/clone
Clone and speak in one request.
Clone and speak in one multipart request. Ten to fifteen seconds of clean, single-speaker audio works best.
| Name | Type | Required | Description |
|---|
text | string | yes | Text to speak in the cloned voice. |
ref_audio | file | yes | Reference clip. Mono WAV recommended; 10–15 s of clean single-speaker audio works best. Max 10 MB. |
ref_text | string | no | Transcript of the clip. Strongly recommended; auto-transcribed when omitted. |
lang | string | no | Language hint for normalization and transcription. |
duration | number | no | Fixed output duration in seconds. Overrides speed. |
speed | number | no | Speed factor. |
guidance_scale | number | no | Classifier-free guidance scale. |
num_step | integer | no | Diffusion steps. |
breaks | string | no | JSON string of auto-pause config, e.g. {"sentence":450,"comma":250}. |
output_format | string | no | wav, mp3, flac or ogg. |
Python
import os, requests
with open("reference.wav", "rb") as fh:
r = requests.post(
"https://voice.saydi.ai/api/tts/clone",
headers={"Authorization": f"Bearer {os.environ['SAYDI_API_KEY']}"},
data={
"text": "Xin chào, đây là giọng đã nhân bản.",
"ref_text": "Đây là câu nói trong clip mẫu.",
"output_format": "mp3",
},
files={"ref_audio": ("reference.wav", fh, "audio/wav")},
timeout=180,
)
r.raise_for_status()
open("cloned.mp3", "wb").write(r.content)
Constraints — Reference audio max 10 MB. Only clone voices you have permission to use; re-uploading the clip every call is slower than a saved voice.
POST /voice/transcribe
Transcribe a reference clip for better cloning.
Transcribe a short reference clip to get ref_text for better cloning accuracy.
| Name | Type | Required | Description |
|---|
audio | file | yes | The reference clip. Mono WAV recommended, max 10 MB. |
lang | string | no | Optional language hint (ISO 639-1). |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/voice/transcribe" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "audio=@reference.wav" \
-F "lang=vi"
200 — Transcription plus a quality verdict. Check ok, not just transcript.
{
"ok": true,
"reason": "ok",
"transcript": "Đây là câu nói trong clip mẫu.",
"word_count": 8,
"char_count": 25
}
Credits & quota
How text and cloning are billed.
Text to speech and cloning are billed by character from one shared quota for studio and API.
| Event | Meter | Charged |
|---|
| POST /v1/audio/speech | characters | input length |
| POST /tts, /tts/file | characters | text length |
| POST /tts/clone | characters | text length |
| POST /voice/transcribe | feature gate | No |
| GET /samples, /v1/models | none | No |
Current plan and usage preflight
# Current subscription (plan, cycle, period end)
curl -sS "https://voice.saydi.ai/api/billing/me" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
# Would this text fit? Cheap check before spending GPU time.
curl -sS "https://voice.saydi.ai/api/billing/quota/preflight?chars=1200" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
# Everything the developer console shows, in one call (see Usage API below):
curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
Usage API
Query balances, usage and call history.
Requires a signed-in user token. Use these endpoints for your own dashboards and alerts.
GET /usage/summary — Credits and request counters
Credit balance (shared with the studio) plus API request counts and a per-day series.
| Name | Type | Required | Description |
|---|
days | integer | no | Series window, 1–365. Default 30. |
cURL
curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Credits come from the shared pool; requests/characters count API traffic only.
{
"ok": true,
"plan_code": "pro",
"period_start": 1758844800.0,
"period_end": 1761436800.0,
"credits": {
"status": "ready",
"unlimited": false,
"limit": 600000,
"used": 41250,
"remaining": 558750,
"extra": 0,
"day_limit": null,
"day_used": 3200,
"day_remaining": null
},
"api": {
"limits": {"api_concurrency": 5, "clone_slots": 50, "max_chars_request": 100000},
"requests": {"total": 412, "success": 401, "errors": 11},
"characters": 41250,
"last_call_at": 1760001234.5,
"status": "ready",
"days": [
{"day": "2026-09-25", "requests": 18, "characters": 2100, "errors": 0},
{"day": "2026-09-26", "requests": 24, "characters": 3200, "errors": 2}
]
}
}
GET /usage/calls — Call history
Newest calls first, including rejected ones, with a cursor for paging.
| Name | Type | Required | Description |
|---|
limit | integer | no | Page size, 1–200. Default 50. |
before | number | no | Cursor: the next_before value from the previous page. |
cURL
curl -sS "https://voice.saydi.ai/api/usage/calls?limit=25" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Rows are retained for 400 days. next_before is null on the last page.
{
"ok": true,
"status": "ready",
"next_before": 1760001000.0,
"calls": [
{
"ts": 1760001234.5,
"method": "POST",
"endpoint": "/v1/audio/speech",
"status": 200,
"chars": 41,
"voice": "vi-adam",
"latency_ms": 3820,
"key_id": "key_1f0c9a2b7d3e4f56",
"limit_reason": ""
},
{
"ts": 1760001000.0,
"method": "POST",
"endpoint": "/tts",
"status": 429,
"chars": 0,
"voice": "",
"latency_ms": 6,
"key_id": "key_1f0c9a2b7d3e4f56",
"limit_reason": "api-key-window"
}
]
}
Rate limits & concurrency
Per-key request and concurrency limits.
Limits are applied per API key; concurrency is per plan.
| Limit | Default | Response when exceeded |
|---|
| Requests per key | 60 / minute | 429, Retry-After, X-OV-Limit: api-key-window |
| Concurrent generations | Pro 5 · Business 10 (trial 2) | 429, Retry-After, X-OV-Limit: api-concurrency |
Retry with backoff
import time, requests
def speak(text, voice, tries=4):
for attempt in range(tries):
r = requests.post(
"https://voice.saydi.ai/api/v1/audio/speech",
headers={"Authorization": f"Bearer {KEY}"},
json={"input": text, "voice": voice},
timeout=180,
)
if r.status_code != 429:
r.raise_for_status()
return r.content
wait = int(r.headers.get("Retry-After", "5"))
time.sleep(max(1, wait) * (2 ** attempt)) # exponential backoff
raise RuntimeError("rate limited after retries")
On 429 — Wait for Retry-After, then retry with backoff and jitter.
Error codes
Error shapes, common codes and retry guidance.
Native endpoints return {"detail": ...}. POST /v1/audio/speech and GET /v1/models return the OpenAI {"error": {...}} shape for EVERY failure — bad credentials (401), a bad field (422), a quota or rate limit (429), a server error (500) — with the response headers (Retry-After, X-OV-Limit) preserved. One deliberate difference: a bad field keeps status 422 here, while OpenAI answers 400 for the same mistake.
Native
{
"detail": "Field 'text' must not be empty."
}
OpenAI-compatible
{
"error": {
"message": "Field 'input' must not be empty.",
"type": "invalid_request_error",
"param": "input",
"code": "invalid_value"
}
}
| Status | Meaning | Retry? |
|---|
| 400 | Malformed request, empty text, or undecodable reference audio | No — fix the request |
| 401 | Missing, unknown, revoked or expired credential | No — fix the credential |
| 403 | Account or token blocked for suspected abuse | No — contact support |
| 404 | Endpoint or voice not found | No |
| 413 | Text too long, or upload above 10 MB | No — split or compress |
| 422 | Request failed schema validation (wrong type or range) | No — check parameter types |
| 429 | Quota exhausted, per-key rate limit, or concurrency cap (on /v1/*, error.code separates them: insufficient_quota vs rate_limit_exceeded) | Yes — after Retry-After |
| 500 | Unexpected server error | Yes — with backoff |
| 503 | Backend temporarily unavailable or not configured | Yes — with backoff |
X-OV-Limit on 403/429 names the limit that blocked the request.
| X-OV-Limit | Status | Cause |
|---|
| api-key-window | 429 | Per-key request rate exceeded |
| api-concurrency | 429 | Too many in-flight generations for the plan |
| usage-limit | 429 / 403 | Credit quota exhausted, or account soft-limited |
| blocked | 403 | Account or token blocked |
| ip-window | 429 | Edge per-IP limit (shared infrastructure, not your key) |
| concurrency | 503 | Server-wide in-flight ceiling reached |
Changelog
Recent API changes and versioning.
| Date | Version | Change | Breaking? |
|---|
| 2026-09-26 | 1.0 | Giai đoạn 1: per-account API keys, per-key rate limit, per-plan concurrency, shared character credits. | — |
Versioning policy — Additive changes ship without notice. Breaking changes get a new /vN prefix and an entry here.
Support
Where to get help.
| Topic | Channel |
|---|
| Integration questions | contact@saydi.ai |
| Higher limits / Enterprise | contact@saydi.ai |
| Suspected abuse or a leaked key | Revoke the key in the dashboard, then contact support |
| Service status | GET /health (no auth required) |