# SaydiVoice API — full reference for AI assistants Version 1.0 · Giai đoạn 1 · updated 2026-09-26 · base URL https://voice.saydi.ai/api This file is generated from the same source as the human documentation at /docs. Prefer it over scraping the HTML. # SaydiVoice API > Text to speech and voice cloning over REST. - Version: 1.0 · Giai đoạn 1 - Updated: 2026-09-26 - Base URL: https://voice.saydi.ai/api - Auth: Authorization: Bearer sv_live_... --- # Introduction Text to speech and voice cloning over one REST API. SaydiVoice turns text into speech and clones a voice from a short reference clip. One bearer token, audio bytes out. > **Note: What you can call** > Text to speech (POST /v1/audio/speech, OpenAI-compatible), voice cloning (POST /tts/clone), transcription (POST /voice/transcribe), plus the /samples catalogue. --- # Quick Start Create a key and get your first audio file. 1. **Create an API key** — Open [/developers/#/keys](/developers/#/keys), create a key and copy the secret (shown once). 2. **Store it safely** — Put it in an environment variable. Never commit it or ship it in client-side code. 3. **Make your first request** — POST to /v1/audio/speech and write the response body to a file. **cURL** ```bash export SAYDI_API_KEY="sv_live_REPLACE_WITH_YOUR_KEY" curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"tts-1","voice":"vi-adam","input":"Xin chào, đây là API của SaydiVoice."}' \ --output hello.mp3 # Inspect what you were charged (characters) and how long it took: # -D - prints the response headers, e.g. X-OV-Chars: 41 ``` **Python** ```python import os, requests API_KEY = os.environ["SAYDI_API_KEY"] BASE = "https://voice.saydi.ai/api" r = requests.post( f"{BASE}/v1/audio/speech", headers={"Authorization": f"Bearer {API_KEY}"}, json={ "model": "tts-1", "voice": "vi-adam", "input": "Xin chào, đây là API của SaydiVoice.", }, timeout=120, ) r.raise_for_status() open("hello.mp3", "wb").write(r.content) print("chars charged:", r.headers.get("X-OV-Chars")) ``` **Node.js** ```javascript import { writeFile } from "node:fs/promises"; const API_KEY = process.env.SAYDI_API_KEY; const BASE = "https://voice.saydi.ai/api"; const res = await fetch(`${BASE}/v1/audio/speech`, { method: "POST", headers: { Authorization: `Bearer ${API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ model: "tts-1", voice: "vi-adam", input: "Xin chào, đây là API của SaydiVoice.", }), }); if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`); await writeFile("hello.mp3", Buffer.from(await res.arrayBuffer())); console.log("chars charged:", res.headers.get("X-OV-Chars")); ``` > **Tip: Expect audio bytes** > Success returns audio bytes (audio/mpeg or audio/wav). A JSON response means an error. --- # Authentication Authenticate with one API key. Send the key as `Authorization: Bearer sv_live_...`. Never put it in URLs, logs or client-side code. **Request headers** ```http Authorization: Bearer sv_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx Content-Type: application/json ``` --- # Manage keys Create, list, rotate and revoke API keys. These routes manage keys and require a signed-in user token, not an API key. ## GET /api-keys — List keys Returns every key on the account. Secrets and hashes are never included. **cURL** ```bash curl -sS "https://voice.saydi.ai/api/api-keys" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" ``` **Response 200** — The account's keys, newest first. ```json { "ok": true, "max_keys": 20, "keys": [ { "id": "key_1f0c9a2b7d3e4f56", "name": "Production", "prefix": "sv_live_Ab12Cd", "scopes": ["tts"], "status": "active", "created_at": 1758844800.0, "last_used_at": 1758848400.0, "expires_at": null, "revoked_at": null } ] } ``` ## POST /api-keys — Create a key Creates a key and returns the secret exactly once, in the key field. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `name` | string | no | A label to recognise the key later. Defaults to "Default key". | | `scopes` | string[] | no | Reserved for later releases. Defaults to ["tts"]. | **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/api-keys" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" \ -H "Content-Type: application/json" \ -d '{"name":"Production"}' ``` **Response 201** — Created. Store key now — it cannot be read again. ```json { "ok": true, "id": "key_1f0c9a2b7d3e4f56", "name": "Production", "prefix": "sv_live_Ab12Cd", "scopes": ["tts"], "status": "active", "created_at": 1758844800.0, "key": "sv_live_Ab12CdEf34Gh56Ij78Kl90Mn12Op34Qr" } ``` **Response 409** — Per-account key limit reached. Revoke an unused key first. ```json {"detail": "API key limit reached (20). Revoke an unused key first."} ``` ## DELETE /api-keys/{key_id} — Revoke a key Revokes immediately. The key stops authenticating on the next request. **cURL** ```bash curl -sS -X DELETE "https://voice.saydi.ai/api/api-keys/key_1f0c9a2b7d3e4f56" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" ``` **Response 200** — Revoked. ```json {"ok": true, "id": "key_1f0c9a2b7d3e4f56", "status": "revoked"} ``` **Response 404** — No such key on this account. ```json {"detail": "API key not found."} ``` ## POST /api-keys/{key_id}/rotate — Rotate a key Issues a new secret for the same key id. The previous secret stops working immediately. Keeps the key id and name, replaces only the secret. **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/api-keys/KEY_ID/rotate" \ -H "Authorization: Bearer $ACCESS_TOKEN" \ -H "Content-Type: application/json" ``` ## GET /api-keys/status — Key management availability Public readiness probe: whether key management is deployed, the per-account cap, and the known scopes. **cURL** ```bash curl -sS "https://voice.saydi.ai/api/api-keys/status" ``` **Response 200** — Readiness snapshot. ```json {"available": true, "max_keys": 20, "scopes": ["tts", "voices"]} ``` --- # POST /v1/audio/speech OpenAI-compatible text to speech. OpenAI-compatible endpoint: same path, same body fields, same audio-bytes response, same error envelope. Read the table below before swapping `base_url` — two fields behave differently on purpose. | Part | OpenAI | SaydiVoice | | --- | --- | --- | | Path, auth, JSON body, audio bytes | POST /v1/audio/speech | Identical | | model | tts-1, tts-1-hd, gpt-4o-mini-tts | tts-1 → 16 steps, tts-1-hd → 32 steps; any other id is accepted and uses the server default | | voice | 13 built-in voices, or {id} for a custom voice | A SaydiVoice voice. The 13 OpenAI names are accepted but ALL resolve to one default (see the callout below); the {id} object form is rejected | | instructions | Free-form style sentence on a fixed voice (gpt-4o-mini-tts) | Voice-design mode with a PRESET vocabulary (gender, age, pitch, accent, whisper). A free-form sentence is 400 with error.param "instructions"; so is combining it with a SaydiVoice voice | | response_format | mp3, opus, aac, flac, wav, pcm | Same six; aac is served as mp3 | | stream_format | audio (chunked), sse (audio events) | audio is accepted but the file arrives in one response; sse is rejected with 400 | | Errors | {"error": {...}}, 400 for a bad field | Same envelope here; a bad field keeps status 422. Native endpoints keep {"detail": ...} | **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `input` | string | yes | Text to speak. Schema maximum 4096 characters; your plan's per-request character cap also applies (Free: 3000), and exceeding it is a 429 with error.code "insufficient_quota". Use POST /tts for longer text. | | `voice` | string | yes | A voice_id from GET /samples (legacy name still accepted), a registered clone voice from POST /v1/voices (`vc_...`), or an OpenAI built-in name (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar). An OpenAI name does NOT select that voice: none exist here, so they all resolve to the alias default (vi-adam unless OMNIVOICE_OPENAI_VOICE_DEFAULT says otherwise) and the response shows X-OV-Voice-Source: openai-alias-default. Any other unknown name falls back to auto-voice. X-OV-Voice always reports the voice that actually spoke, or empty for auto. | | `model` | string | no | "tts-1" = 16 diffusion steps (faster), "tts-1-hd" = 32 steps (higher quality). Omitted = server default. | | `response_format` | string | no | mp3 (default), opus, aac, flac, wav, pcm. aac is currently served as mp3. | | `speed` | number | no | Speech speed, 0.25–4.0. Default 1.0. | | `instructions` | string | no | Style instruction, mapped to this server's voice-design prompt (same engine as `instruct` on POST /tts). The vocabulary is PRESET — gender, age, pitch, accent, whisper — for example "female, british accent, whisper". A free-form style sentence (what OpenAI's gpt-4o-mini-tts accepts, e.g. "Speak in a cheerful and positive tone.") is refused with 400, error.param "instructions", and the message lists the valid items. It also cannot be combined with a SaydiVoice voice: design mode builds the speaker from the instruction and has no reference clip. | | `stream_format` | string | no | Compatibility field. 'audio' is accepted and the whole file arrives in one response; 'sse' is rejected with 400 because audio events are not implemented. | > **Warning: OpenAI voice names are accepted, but they do not select a voice** > None of alloy/ash/ballad/coral/echo/fable/nova/onyx/sage/shimmer/verse/marin/cedar exist in this catalogue, so every one of them resolves to a single default voice (vi-adam unless OMNIVOICE_OPENAI_VOICE_DEFAULT says otherwise). The response says so: `X-OV-Voice-Source: openai-alias-default` plus `X-OV-Voice-Hint`. To choose a voice, pass a `voice_id` from GET /samples — X-OV-Voice then reports the voice that actually spoke, and `saydivoice` as the source. **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-1-hd", "voice": "vi-adam", "input": "Higher quality, 32 diffusion steps.", "response_format": "wav", "speed": 1.0 }' \ --output speech.wav ``` **OpenAI SDK (base_url swap)** ```python from openai import OpenAI client = OpenAI( api_key=os.environ["SAYDI_API_KEY"], base_url="https://voice.saydi.ai/api/v1", # only this line changes ) # The body is unchanged. `voice` is a SaydiVoice voice_id from GET /samples: # OpenAI's built-in names are accepted but all resolve to one default voice. with client.audio.speech.with_streaming_response.create( model="tts-1", voice="vi-adam", input="Drop-in replacement.", ) as response: response.stream_to_file("speech.mp3") ``` > **Note: Metering headers** > Each audio response reports cost and timing: X-OV-Chars, X-Audio-Duration, X-Generation-Time, X-RTF. | Header | Meaning | | --- | --- | | X-OV-Chars | Characters charged for this request | | X-OV-Voice | Voice actually used (after fallback) | | X-Audio-Duration | Generated audio length in seconds | | X-Generation-Time | Compute time in seconds | | X-RTF | Real-time factor: generation time / audio duration | --- # POST /tts Full-control text to speech. Use this endpoint for long text, fixed duration, auto pauses, voice design or sampling controls. ## Voice selection — three modes | Mode | Set this | Result | | --- | --- | --- | | Voice cloning | sample (+ optional ref_text) | Speaks in a catalogue voice copied from its reference clip | | Voice design | instruct | The model creates a voice matching a text description | | Auto voice | neither | The model picks a voice for the text | **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `text` | string | yes | Text to synthesize. | | `sample` | string | no | Voice_id from GET /samples (legacy name still accepted). Enables cloning mode. | | `voice_id` | string | no | A registered clone voice from POST /v1/voices (e.g. `vc_9f2c...`). No reference audio needed; the server uses the stored clip and transcript. | | `instruct` | string | no | Voice design description, e.g. "female, low pitch, british accent". Ignored when sample is set. | | `ref_text` | string | no | Override the reference transcript for the sample. Improves cloning accuracy. | | `lang` | string | no | Language hint (ISO 639-1). Used for text normalization and pronunciation. | | `breaks` | object | no | Auto pauses in ms per punctuation type: {"sentence":450,"comma":250,"semicolon":300,"paragraph":600}. | | `duration` | number | no | Fixed output duration in seconds (max 600). Overrides speed. Useful to fit a time slot. | | `speed` | number | no | Speed factor, 0–10. Ignored when duration is set. | | `num_step` | integer | no | Diffusion steps 1–256. Default 32; 16 is faster. | | `guidance_scale` | number | no | Classifier-free guidance, 0–50. Default 2.0. | | `t_shift` | number | no | Noise-schedule shift, 0–1. Default 0.1. | | `denoise` | boolean | no | Prepend the denoising token. Default true. | | `preprocess_prompt` | boolean | no | Clean the reference clip before cloning. Default true. | | `postprocess_output` | boolean | no | Trim long silences from the output. Default true. | | `audio_chunk_duration` | number | no | Target chunk length in seconds for long text (max 120). Default 15. | | `audio_chunk_threshold` | number | no | Estimated duration that triggers chunking (max 600). Default 30. | | `output_format` | string | no | wav, mp3, flac or ogg. Default from server config. | **Long text with pauses and a fixed duration** ```bash curl -sS -X POST "https://voice.saydi.ai/api/tts" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "Câu thứ nhất. Câu thứ hai, có dấu phẩy.", "sample": "vi-adam", "lang": "vi", "breaks": {"sentence": 450, "comma": 250}, "num_step": 32, "output_format": "mp3" }' \ --output long.mp3 ``` **Voice design (no reference clip)** ```bash curl -sS -X POST "https://voice.saydi.ai/api/tts" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "A calm narrator for a documentary.", "instruct": "male, middle-aged, low pitch, american accent", "lang": "en" }' \ --output designed.mp3 ``` --- # POST /tts/file Generate speech from a .txt file. Send the script as a UTF-8 .txt file instead of embedding a long string in JSON. **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/tts/file" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -F "text_file=@script.txt" \ -F "sample=vi-adam" \ -F "lang=vi" \ -F "output_format=mp3" \ --output script.mp3 ``` Accepts every field of POST /tts; use `text_file` instead of `text`. --- # GET /samples List voices and metadata. Pass `voice_id` as `voice` (OpenAI-compatible) or `sample` (native). Legacy `name` still works; use `display_name` for UI labels. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `lang` | string | no | Return only voices whose primary language matches, e.g. "vi". | | `limit` | integer | no | Return at most this many voices, after sorting by usage. | **cURL** ```bash curl -sS "https://voice.saydi.ai/api/samples?lang=vi&limit=5" \ -H "Authorization: Bearer $SAYDI_API_KEY" ``` **Response 200** — An array of voice objects. ```json [ { "name": "adam-11labs-vi", "voice_id": "vi-adam", "audio_file": "adam-11labs-vi.wav", "has_transcript": true, "transcript_preview": "Xin chào, tôi là Adam...", "display_name": "Adam — Giọng hot tiktok", "gender": "male", "age": "young", "accent": "vietnamese", "locale": "vi-VN", "language": "vi", "languages": ["vi"], "use_case": "conversational", "category": "high_quality", "featured": true, "usage_count": 12840 } ] ``` ## GET /samples/{name}/audio — Preview a voice Streams the reference clip for a voice, so you can audition it before generating. **cURL** ```bash curl -sS "https://voice.saydi.ai/api/samples/vi-adam/audio" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ --output preview.wav ``` ## GET /v1/models — List models OpenAI-compatible model list: tts-1 (16 steps) and tts-1-hd (32 steps). **cURL** ```bash curl -sS "https://voice.saydi.ai/api/v1/models" \ -H "Authorization: Bearer $SAYDI_API_KEY" ``` --- # POST /v1/voices Register a voice once, reuse it by voice_id. Upload a clip and a name. SaydiVoice checks the audio, picks its best continuous stretch of speech, transcribes it, and removes background music when it finds any — then returns a stable `voice_id`. Use that id in `POST /v1/audio/speech` or `POST /tts`: no clip to re-upload, and the reference is cached server-side so later generations start faster. Voices registered here also appear in Studio → My Voices, and voices you create in Studio appear here. ## POST /v1/voices — Register a voice Multipart upload. Returns the voice object; `status` is `ready` when it is usable right away. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `file` | file | yes | Reference clip: WAV, MP3, M4A, OGG, FLAC or WebM. Any length up to 5 minutes; the best stretch is selected for you. | | `name` | string | yes | Display name. Not unique — identify the voice by `id`, never by name. | | `consent` | string | yes | Confirm you have the right to clone this voice. Send `true`, or JSON `{"accepted": true}`. Recorded with your IP and the terms version. | | `language` | string | no | ISO-639-1 hint (e.g. `vi`). Also picks the language of the preview sentence. | | `ref_text` | string | no | Transcript of the clip. Omit it and we transcribe for you; send it (when your clip is already prepared) to skip that step. | | `separate` | string | no | `auto` (default) separates background music only when detected; `true` always does; `false` never does. | | `Idempotency-Key` | string | no | Header. Retrying with the same value returns the voice that already exists instead of creating a second one — so a timeout never costs you a clone slot. | **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/v1/voices" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -F "file=@my-recording.m4a" \ -F "name=Giọng của tôi" \ -F "language=vi" \ -F "consent=true" # -> {"id":"vc_9f2c...","status":"ready","ref_text":"...","quality":{...}} # Then speak with it — no clip to upload: curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"tts-1","voice":"vc_9f2c...","input":"Xin chào từ giọng của tôi."}' \ --output hello.mp3 ``` **Python** ```python import os, requests key = os.environ["SAYDI_API_KEY"] base = "https://voice.saydi.ai/api" with open("my-recording.m4a", "rb") as fh: voice = requests.post( f"{base}/v1/voices", headers={"Authorization": f"Bearer {key}"}, files={"file": ("my-recording.m4a", fh, "audio/mp4")}, data={"name": "Giọng của tôi", "language": "vi", "consent": "true"}, timeout=300, ) voice.raise_for_status() voice = voice.json() print(voice["id"], voice["status"], voice.get("quality", {}).get("warnings")) audio = requests.post( f"{base}/v1/audio/speech", headers={"Authorization": f"Bearer {key}"}, json={"model": "tts-1", "voice": voice["id"], "input": "Xin chào từ giọng của tôi."}, timeout=180, ) audio.raise_for_status() open("hello.mp3", "wb").write(audio.content) ``` **Response 201** — The voice was created. ```json { "id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12", "object": "voice", "name": "Giọng của tôi", "status": "ready", "language": "vi", "ref_seconds": 7.4, "ref_text": "Xin chào, đây là đoạn nói trong clip mẫu.", "quality": { "speech_seconds": 6.9, "snr_db": 26.4, "trimmed_seconds": 42.1, "music_detected": false, "separated": false, "warnings": [] }, "consent": {"terms_version": "voice-terms-2026-01"}, "preview_url": "/v1/voices/vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12/preview" } ``` **Response 200** — An idempotent retry: this voice already existed and is returned unchanged. ```json { "id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12", "status": "ready", "idempotent_replay": true } ``` > **Note: processing vs ready** > `ready` means usable now — that is the normal case for a short clean clip. `processing` means the clip had music under the voice, so separation is still running: poll `GET /v1/voices/{voice_id}` (usually seconds) before synthesising. A `failed` voice frees its slot and carries an `error.code` you can act on. Two errors are worth handling separately: `503 storage_unavailable` is temporary — retry with backoff; `410 ref_missing` means the stored audio is gone for good, so register the voice again (it does not use an extra slot). ## GET /v1/voices — List and manage voices List your voices, fetch one, rename it, delete it, or hear a preview. | Endpoint | What it does | | --- | --- | | `GET /v1/voices` | This account's voices, newest first, with `usage` (slots used/limit). Filter with `?status=ready`. | | `GET /v1/voices/{voice_id}` | One voice: `status`, `ref_text` and the quality report. | | `PATCH /v1/voices/{voice_id}` | Rename: JSON `{"name": "..."}`. The `id` never changes. | | `DELETE /v1/voices/{voice_id}` | Delete the voice and its stored audio, freeing its clone slot. Immediate. | | `GET /v1/voices/{voice_id}/preview` | MP3 of a short sentence in that voice. Rendered once and cached. Not charged against your characters. | **cURL** ```bash curl -sS "https://voice.saydi.ai/api/v1/voices" \ -H "Authorization: Bearer $SAYDI_API_KEY" curl -sS -X PATCH "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Giọng nam trầm"}' curl -sS "https://voice.saydi.ai/api/v1/voices/$VOICE_ID/preview" \ -H "Authorization: Bearer $SAYDI_API_KEY" --output preview.mp3 curl -sS -X DELETE "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \ -H "Authorization: Bearer $SAYDI_API_KEY" ``` > **Warning: Consent and slot limits** > Only register voices you have the right to use; the API refuses a registration without `consent`. Each plan includes a number of clone slots (`clone_slots`) — creating a voice beyond it returns `402` with `code: voice_limit_reached`. A failed registration does not consume a slot. --- # POST /tts/clone Clone and speak in one request. Clone and speak in one multipart request. Ten to fifteen seconds of clean, single-speaker audio works best. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `text` | string | yes | Text to speak in the cloned voice. | | `ref_audio` | file | yes | Reference clip. Mono WAV recommended; 10–15 s of clean single-speaker audio works best. Max 10 MB. | | `ref_text` | string | no | Transcript of the clip. Strongly recommended; auto-transcribed when omitted. | | `lang` | string | no | Language hint for normalization and transcription. | | `duration` | number | no | Fixed output duration in seconds. Overrides speed. | | `speed` | number | no | Speed factor. | | `guidance_scale` | number | no | Classifier-free guidance scale. | | `num_step` | integer | no | Diffusion steps. | | `breaks` | string | no | JSON string of auto-pause config, e.g. {"sentence":450,"comma":250}. | | `output_format` | string | no | wav, mp3, flac or ogg. | **Python** ```python import os, requests with open("reference.wav", "rb") as fh: r = requests.post( "https://voice.saydi.ai/api/tts/clone", headers={"Authorization": f"Bearer {os.environ['SAYDI_API_KEY']}"}, data={ "text": "Xin chào, đây là giọng đã nhân bản.", "ref_text": "Đây là câu nói trong clip mẫu.", "output_format": "mp3", }, files={"ref_audio": ("reference.wav", fh, "audio/wav")}, timeout=180, ) r.raise_for_status() open("cloned.mp3", "wb").write(r.content) ``` > **Warning: Constraints** > Reference audio max 10 MB. Only clone voices you have permission to use; re-uploading the clip every call is slower than a saved voice. --- # POST /voice/transcribe Transcribe a reference clip for better cloning. Transcribe a short reference clip to get `ref_text` for better cloning accuracy. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `audio` | file | yes | The reference clip. Mono WAV recommended, max 10 MB. | | `lang` | string | no | Optional language hint (ISO 639-1). | **cURL** ```bash curl -sS -X POST "https://voice.saydi.ai/api/voice/transcribe" \ -H "Authorization: Bearer $SAYDI_API_KEY" \ -F "audio=@reference.wav" \ -F "lang=vi" ``` **Response 200** — Transcription plus a quality verdict. Check ok, not just transcript. ```json { "ok": true, "reason": "ok", "transcript": "Đây là câu nói trong clip mẫu.", "word_count": 8, "char_count": 25 } ``` --- # Credits & quota How text and cloning are billed. Text to speech and cloning are billed by character from one shared quota for studio and API. | Event | Meter | Charged | | --- | --- | --- | | POST /v1/audio/speech | characters | input length | | POST /tts, /tts/file | characters | text length | | POST /tts/clone | characters | text length | | POST /voice/transcribe | feature gate | No | | GET /samples, /v1/models | none | No | **Current plan and usage preflight** ```bash # Current subscription (plan, cycle, period end) curl -sS "https://voice.saydi.ai/api/billing/me" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" # Would this text fit? Cheap check before spending GPU time. curl -sS "https://voice.saydi.ai/api/billing/quota/preflight?chars=1200" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" # Everything the developer console shows, in one call (see Usage API below): curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" ``` --- # Usage API Query balances, usage and call history. Requires a signed-in user token. Use these endpoints for your own dashboards and alerts. ## GET /usage/summary — Credits and request counters Credit balance (shared with the studio) plus API request counts and a per-day series. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `days` | integer | no | Series window, 1–365. Default 30. | **cURL** ```bash curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" ``` **Response 200** — Credits come from the shared pool; requests/characters count API traffic only. ```json { "ok": true, "plan_code": "pro", "period_start": 1758844800.0, "period_end": 1761436800.0, "credits": { "status": "ready", "unlimited": false, "limit": 600000, "used": 41250, "remaining": 558750, "extra": 0, "day_limit": null, "day_used": 3200, "day_remaining": null }, "api": { "limits": {"api_concurrency": 5, "clone_slots": 50, "max_chars_request": 100000}, "requests": {"total": 412, "success": 401, "errors": 11}, "characters": 41250, "last_call_at": 1760001234.5, "status": "ready", "days": [ {"day": "2026-09-25", "requests": 18, "characters": 2100, "errors": 0}, {"day": "2026-09-26", "requests": 24, "characters": 3200, "errors": 2} ] } } ``` ## GET /usage/calls — Call history Newest calls first, including rejected ones, with a cursor for paging. **Parameters** | Name | Type | Required | Description | | --- | --- | --- | --- | | `limit` | integer | no | Page size, 1–200. Default 50. | | `before` | number | no | Cursor: the next_before value from the previous page. | **cURL** ```bash curl -sS "https://voice.saydi.ai/api/usage/calls?limit=25" \ -H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" ``` **Response 200** — Rows are retained for 400 days. next_before is null on the last page. ```json { "ok": true, "status": "ready", "next_before": 1760001000.0, "calls": [ { "ts": 1760001234.5, "method": "POST", "endpoint": "/v1/audio/speech", "status": 200, "chars": 41, "voice": "vi-adam", "latency_ms": 3820, "key_id": "key_1f0c9a2b7d3e4f56", "limit_reason": "" }, { "ts": 1760001000.0, "method": "POST", "endpoint": "/tts", "status": 429, "chars": 0, "voice": "", "latency_ms": 6, "key_id": "key_1f0c9a2b7d3e4f56", "limit_reason": "api-key-window" } ] } ``` --- # Rate limits & concurrency Per-key request and concurrency limits. Limits are applied per API key; concurrency is per plan. | Limit | Default | Response when exceeded | | --- | --- | --- | | Requests per key | 60 / minute | 429, Retry-After, X-OV-Limit: api-key-window | | Concurrent generations | Pro 5 · Business 10 (trial 2) | 429, Retry-After, X-OV-Limit: api-concurrency | **Retry with backoff** ```python import time, requests def speak(text, voice, tries=4): for attempt in range(tries): r = requests.post( "https://voice.saydi.ai/api/v1/audio/speech", headers={"Authorization": f"Bearer {KEY}"}, json={"input": text, "voice": voice}, timeout=180, ) if r.status_code != 429: r.raise_for_status() return r.content wait = int(r.headers.get("Retry-After", "5")) time.sleep(max(1, wait) * (2 ** attempt)) # exponential backoff raise RuntimeError("rate limited after retries") ``` > **Tip: On 429** > Wait for Retry-After, then retry with backoff and jitter. --- # Error codes Error shapes, common codes and retry guidance. Native endpoints return `{"detail": ...}`. POST /v1/audio/speech and GET /v1/models return the OpenAI `{"error": {...}}` shape for EVERY failure — bad credentials (401), a bad field (422), a quota or rate limit (429), a server error (500) — with the response headers (Retry-After, X-OV-Limit) preserved. One deliberate difference: a bad field keeps status 422 here, while OpenAI answers 400 for the same mistake. **Native** ```json { "detail": "Field 'text' must not be empty." } ``` **OpenAI-compatible** ```json { "error": { "message": "Field 'input' must not be empty.", "type": "invalid_request_error", "param": "input", "code": "invalid_value" } } ``` | Status | Meaning | Retry? | | --- | --- | --- | | 400 | Malformed request, empty text, or undecodable reference audio | No — fix the request | | 401 | Missing, unknown, revoked or expired credential | No — fix the credential | | 403 | Account or token blocked for suspected abuse | No — contact support | | 404 | Endpoint or voice not found | No | | 413 | Text too long, or upload above 10 MB | No — split or compress | | 422 | Request failed schema validation (wrong type or range) | No — check parameter types | | 429 | Quota exhausted, per-key rate limit, or concurrency cap (on /v1/*, error.code separates them: insufficient_quota vs rate_limit_exceeded) | Yes — after Retry-After | | 500 | Unexpected server error | Yes — with backoff | | 503 | Backend temporarily unavailable or not configured | Yes — with backoff | `X-OV-Limit` on 403/429 names the limit that blocked the request. | X-OV-Limit | Status | Cause | | --- | --- | --- | | api-key-window | 429 | Per-key request rate exceeded | | api-concurrency | 429 | Too many in-flight generations for the plan | | usage-limit | 429 / 403 | Credit quota exhausted, or account soft-limited | | blocked | 403 | Account or token blocked | | ip-window | 429 | Edge per-IP limit (shared infrastructure, not your key) | | concurrency | 503 | Server-wide in-flight ceiling reached | --- # Changelog Recent API changes and versioning. | Date | Version | Change | Breaking? | | --- | --- | --- | --- | | 2026-09-26 | 1.0 | Giai đoạn 1: per-account API keys, per-key rate limit, per-plan concurrency, shared character credits. | — | > **Note: Versioning policy** > Additive changes ship without notice. Breaking changes get a new /vN prefix and an entry here. --- # Support Where to get help. | Topic | Channel | | --- | --- | | Integration questions | contact@saydi.ai | | Higher limits / Enterprise | contact@saydi.ai | | Suspected abuse or a leaked key | Revoke the key in the dashboard, then contact support | | Service status | GET /health (no auth required) |