Tài liệu API SaydiVoice
Trang này cần JavaScript để hiển thị dạng tương tác. Toàn bộ nội dung tham chiếu nằm ngay dưới đây, và cũng có sẵn ở dạng văn bản thuần tại /llms-full.txt.
Base URL https://voice.saydi.ai/api. Xác thực bằng Authorization: Bearer sv_live_….
Giới thiệu
Text to speech và nhân bản giọng nói qua một REST API.
SaydiVoice biến văn bản thành giọng nói và nhân bản giọng từ một clip mẫu ngắn. Một bearer token, kết quả là bytes audio.
Bạn gọi được gì — Text to speech (POST /v1/audio/speech, tương thích OpenAI), nhân bản giọng (POST /tts/clone), phiên âm (POST /voice/transcribe), và thư viện /samples.
Bắt đầu nhanh
Tạo key và nhận file audio đầu tiên.
- Tạo API key — Mở /developers/#/keys, tạo key và copy secret (chỉ hiện một lần).
- Lưu an toàn — Lưu trong biến môi trường. Không commit và không để trong code phía client.
- Gọi request đầu tiên — POST tới /v1/audio/speech và ghi response body ra file.
cURL
export SAYDI_API_KEY="sv_live_REPLACE_WITH_YOUR_KEY"
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"vi-adam","input":"Xin chào, đây là API của SaydiVoice."}' \
--output hello.mp3
# Inspect what you were charged (characters) and how long it took:
# -D - prints the response headers, e.g. X-OV-Chars: 41
Python
import os, requests
API_KEY = os.environ["SAYDI_API_KEY"]
BASE = "https://voice.saydi.ai/api"
r = requests.post(
f"{BASE}/v1/audio/speech",
headers={"Authorization": f"Bearer {API_KEY}"},
json={
"model": "tts-1",
"voice": "vi-adam",
"input": "Xin chào, đây là API của SaydiVoice.",
},
timeout=120,
)
r.raise_for_status()
open("hello.mp3", "wb").write(r.content)
print("chars charged:", r.headers.get("X-OV-Chars"))
Node.js
import { writeFile } from "node:fs/promises";
const API_KEY = process.env.SAYDI_API_KEY;
const BASE = "https://voice.saydi.ai/api";
const res = await fetch(`${BASE}/v1/audio/speech`, {
method: "POST",
headers: {
Authorization: `Bearer ${API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "tts-1",
voice: "vi-adam",
input: "Xin chào, đây là API của SaydiVoice.",
}),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
await writeFile("hello.mp3", Buffer.from(await res.arrayBuffer()));
console.log("chars charged:", res.headers.get("X-OV-Chars"));
Kết quả là audio bytes — Thành công trả audio bytes (audio/mpeg hoặc audio/wav). Response JSON là lỗi.
Xác thực
Xác thực bằng một API key.
Gửi key qua header Authorization: Bearer sv_live_.... Không để key trong URL, log hay code phía client.
Request headers
Authorization: Bearer sv_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Content-Type: application/json
Quản lý key
Tạo, liệt kê, rotate và thu hồi API key.
Các route này quản lý key và cần access token của người dùng, không phải API key.
GET /api-keys — Liệt kê key
Trả về toàn bộ key của tài khoản. Không bao giờ trả secret hay hash.
cURL
curl -sS "https://voice.saydi.ai/api/api-keys" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Danh sách key, mới nhất trước.
{
"ok": true,
"max_keys": 20,
"keys": [
{
"id": "key_1f0c9a2b7d3e4f56",
"name": "Production",
"prefix": "sv_live_Ab12Cd",
"scopes": ["tts"],
"status": "active",
"created_at": 1758844800.0,
"last_used_at": 1758848400.0,
"expires_at": null,
"revoked_at": null
}
]
}
POST /api-keys — Tạo key
Tạo key và trả secret đúng một lần, ở trường key.
| Name | Type | Required | Description |
|---|
name | string | no | Nhãn để nhận biết key. Mặc định "Default key". |
scopes | string[] | no | Dành cho bản sau. Mặc định ["tts"]. |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/api-keys" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"Production"}'
201 — Đã tạo. Lưu key ngay — không thể đọc lại.
{
"ok": true,
"id": "key_1f0c9a2b7d3e4f56",
"name": "Production",
"prefix": "sv_live_Ab12Cd",
"scopes": ["tts"],
"status": "active",
"created_at": 1758844800.0,
"key": "sv_live_Ab12CdEf34Gh56Ij78Kl90Mn12Op34Qr"
}
409 — Đã đạt giới hạn số key của tài khoản. Thu hồi key không dùng trước.
{"detail": "API key limit reached (20). Revoke an unused key first."}
DELETE /api-keys/{key_id} — Thu hồi key
Thu hồi ngay. Key ngừng xác thực ở request kế tiếp.
cURL
curl -sS -X DELETE "https://voice.saydi.ai/api/api-keys/key_1f0c9a2b7d3e4f56" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Đã thu hồi.
{"ok": true, "id": "key_1f0c9a2b7d3e4f56", "status": "revoked"}
404 — Không có key này trên tài khoản.
{"detail": "API key not found."}
POST /api-keys/{key_id}/rotate — Rotate key
Cấp secret mới cho cùng key id. Secret cũ ngừng hoạt động ngay.
Giữ nguyên id và tên key, chỉ thay secret.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/api-keys/KEY_ID/rotate" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json"
GET /api-keys/status — Trạng thái quản lý key
Probe công khai: quản lý key có sẵn sàng không, giới hạn key mỗi tài khoản, và các scope.
cURL
curl -sS "https://voice.saydi.ai/api/api-keys/status"
200 — Ảnh chụp trạng thái.
{"available": true, "max_keys": 20, "scopes": ["tts", "voices"]}
POST /v1/audio/speech
Text to speech tương thích OpenAI.
Endpoint tương thích OpenAI: cùng path, cùng field body, cùng dạng response bytes audio, cùng error envelope. Đọc bảng dưới trước khi đổi base_url — có hai field hành xử khác một cách có chủ ý.
| Phần | OpenAI | SaydiVoice |
|---|
| Path, auth, JSON body, audio bytes | POST /v1/audio/speech | Identical |
| model | tts-1, tts-1-hd, gpt-4o-mini-tts | tts-1 → 16 steps, tts-1-hd → 32 steps; any other id is accepted and uses the server default |
| voice | 13 built-in voices, or {id} for a custom voice | A SaydiVoice voice. The 13 OpenAI names are accepted but ALL resolve to one default (see the callout below); the {id} object form is rejected |
| instructions | Free-form style sentence on a fixed voice (gpt-4o-mini-tts) | Voice-design mode with a PRESET vocabulary (gender, age, pitch, accent, whisper). A free-form sentence is 400 with error.param "instructions"; so is combining it with a SaydiVoice voice |
| response_format | mp3, opus, aac, flac, wav, pcm | Same six; aac is served as mp3 |
| stream_format | audio (chunked), sse (audio events) | audio is accepted but the file arrives in one response; sse is rejected with 400 |
| Errors | {"error": {...}}, 400 for a bad field | Same envelope here; a bad field keeps status 422. Native endpoints keep {"detail": ...} |
| Name | Type | Required | Description |
|---|
input | string | yes | Văn bản cần đọc. Schema tối đa 4096 ký tự; trần ký tự mỗi lượt của gói cũng áp dụng (Free: 3000), vượt là 429 với error.code "insufficient_quota". Văn bản dài hơn dùng POST /tts. |
voice | string | yes | voice_id từ GET /samples (name cũ vẫn được chấp nhận), giọng đã đăng ký từ POST /v1/voices (vc_...), hoặc tên giọng OpenAI built-in (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar). Tên OpenAI KHÔNG chọn được giọng đó: catalogue không có giọng nào trùng tên, nên tất cả về giọng mặc định (vi-adam, đổi bằng OMNIVOICE_OPENAI_VOICE_DEFAULT) và response ghi rõ X-OV-Voice-Source: openai-alias-default. Tên lạ khác thì tự chọn giọng (auto). X-OV-Voice luôn báo giọng thực sự đã đọc, hoặc rỗng khi auto. |
model | string | no | "tts-1" = 16 bước (nhanh hơn), "tts-1-hd" = 32 bước (chất lượng cao hơn). Bỏ trống = mặc định server. |
response_format | string | no | mp3 (mặc định), opus, aac, flac, wav, pcm. Hiện aac được trả về dạng mp3. |
speed | number | no | Tốc độ đọc, 0.25–4.0. Mặc định 1.0. |
instructions | string | no | Câu lệnh style, map sang prompt voice-design của server này (cùng engine với instruct ở POST /tts). Từ vựng là PRESET — gender, age, pitch, accent, whisper — ví dụ "female, british accent, whisper". Một câu style tự do (kiểu OpenAI gpt-4o-mini-tts, ví dụ "Speak in a cheerful and positive tone.") bị từ chối với 400, error.param là "instructions", và message liệt kê các từ hợp lệ. Cũng không kết hợp được với giọng SaydiVoice: design mode dựng người nói từ câu lệnh và không có clip tham chiếu. |
stream_format | string | no | Field tương thích. 'audio' được nhận và cả file về trong một response; 'sse' bị từ chối với 400 vì audio events chưa có. |
Tên giọng OpenAI được nhận, nhưng không chọn được giọng — Catalogue không có giọng nào tên alloy/ash/ballad/coral/echo/fable/nova/onyx/sage/shimmer/verse/marin/cedar, nên mọi tên trong số đó đều về một giọng mặc định (vi-adam, đổi bằng OMNIVOICE_OPENAI_VOICE_DEFAULT). Response ghi rõ: X-OV-Voice-Source: openai-alias-default kèm X-OV-Voice-Hint. Muốn chọn giọng, truyền voice_id từ GET /samples — khi đó X-OV-Voice báo đúng giọng đã đọc và source là saydivoice.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1-hd",
"voice": "vi-adam",
"input": "Higher quality, 32 diffusion steps.",
"response_format": "wav",
"speed": 1.0
}' \
--output speech.wav
OpenAI SDK (base_url swap)
from openai import OpenAI
client = OpenAI(
api_key=os.environ["SAYDI_API_KEY"],
base_url="https://voice.saydi.ai/api/v1", # only this line changes
)
# The body is unchanged. `voice` is a SaydiVoice voice_id from GET /samples:
# OpenAI's built-in names are accepted but all resolve to one default voice.
with client.audio.speech.with_streaming_response.create(
model="tts-1",
voice="vi-adam",
input="Drop-in replacement.",
) as response:
response.stream_to_file("speech.mp3")
Header tính phí — Mỗi response audio báo chi phí và thời gian: X-OV-Chars, X-Audio-Duration, X-Generation-Time, X-RTF.
| Header | Ý nghĩa |
|---|
| X-OV-Chars | Characters charged for this request |
| X-OV-Voice | Voice actually used (after fallback) |
| X-Audio-Duration | Generated audio length in seconds |
| X-Generation-Time | Compute time in seconds |
| X-RTF | Real-time factor: generation time / audio duration |
POST /tts
Text to speech đầy đủ tham số.
Dùng endpoint này cho văn bản dài, ép thời lượng, tự động ngắt nghỉ, voice design hoặc sampling.
Chọn giọng — ba chế độ
| Chế độ | Truyền gì | Kết quả |
|---|
| Voice cloning | sample (+ optional ref_text) | Speaks in a catalogue voice copied from its reference clip |
| Voice design | instruct | The model creates a voice matching a text description |
| Auto voice | neither | The model picks a voice for the text |
| Name | Type | Required | Description |
|---|
text | string | yes | Văn bản cần đọc. |
sample | string | no | voice_id từ GET /samples (name cũ vẫn được chấp nhận). Bật chế độ nhân bản. |
voice_id | string | no | Giọng đã đăng ký từ POST /v1/voices (ví dụ vc_9f2c...). Không cần gửi audio mẫu; server dùng clip và transcript đã lưu. |
instruct | string | no | Mô tả giọng, ví dụ "female, low pitch, british accent". Bỏ qua nếu có sample. |
ref_text | string | no | Ghi đè transcript của clip mẫu. Giúp nhân bản chính xác hơn. |
lang | string | no | Gợi ý ngôn ngữ (ISO 639-1). Dùng để chuẩn hoá văn bản và phát âm. |
breaks | object | no | Tự chèn khoảng nghỉ (ms) theo loại dấu câu: {"sentence":450,"comma":250,"semicolon":300,"paragraph":600}. |
duration | number | no | Ép độ dài output, giây (tối đa 600). Ghi đè speed. Hữu ích để khớp khung thời gian. |
speed | number | no | Hệ số tốc độ, 0–10. Bỏ qua nếu có duration. |
num_step | integer | no | Số bước diffusion 1–256. Mặc định 32; 16 nhanh hơn. |
guidance_scale | number | no | Classifier-free guidance, 0–50. Mặc định 2.0. |
t_shift | number | no | Dịch noise schedule, 0–1. Mặc định 0.1. |
denoise | boolean | no | Thêm token denoise ở đầu. Mặc định true. |
preprocess_prompt | boolean | no | Làm sạch clip mẫu trước khi nhân bản. Mặc định true. |
postprocess_output | boolean | no | Cắt khoảng lặng dài trong output. Mặc định true. |
audio_chunk_duration | number | no | Độ dài mỗi chunk cho văn bản dài, giây (tối đa 120). Mặc định 15. |
audio_chunk_threshold | number | no | Độ dài ước tính để bắt đầu chia chunk (tối đa 600). Mặc định 30. |
output_format | string | no | wav, mp3, flac hoặc ogg. Mặc định theo cấu hình server. |
Long text with pauses and a fixed duration
curl -sS -X POST "https://voice.saydi.ai/api/tts" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Câu thứ nhất. Câu thứ hai, có dấu phẩy.",
"sample": "vi-adam",
"lang": "vi",
"breaks": {"sentence": 450, "comma": 250},
"num_step": 32,
"output_format": "mp3"
}' \
--output long.mp3
Voice design (no reference clip)
curl -sS -X POST "https://voice.saydi.ai/api/tts" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "A calm narrator for a documentary.",
"instruct": "male, middle-aged, low pitch, american accent",
"lang": "en"
}' \
--output designed.mp3
POST /tts/file
Tạo giọng nói từ file .txt.
Gửi kịch bản dưới dạng file .txt UTF-8 thay vì nhúng chuỗi dài vào JSON.
cURL
curl -sS -X POST "https://voice.saydi.ai/api/tts/file" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "text_file=@script.txt" \
-F "sample=vi-adam" \
-F "lang=vi" \
-F "output_format=mp3" \
--output script.mp3
Nhận mọi field của POST /tts; dùng text_file thay cho text.
GET /samples
Liệt kê giọng đọc và metadata.
Truyền voice_id vào voice (OpenAI-compatible) hoặc sample (native). name cũ vẫn dùng được; dùng display_name để hiển thị.
| Name | Type | Required | Description |
|---|
lang | string | no | Chỉ trả giọng có ngôn ngữ chính khớp, ví dụ "vi". |
limit | integer | no | Trả tối đa bao nhiêu giọng, sau khi sắp theo mức dùng. |
cURL
curl -sS "https://voice.saydi.ai/api/samples?lang=vi&limit=5" \
-H "Authorization: Bearer $SAYDI_API_KEY"
200 — Mảng các object giọng nói.
[
{
"name": "adam-11labs-vi",
"voice_id": "vi-adam",
"audio_file": "adam-11labs-vi.wav",
"has_transcript": true,
"transcript_preview": "Xin chào, tôi là Adam...",
"display_name": "Adam — Giọng hot tiktok",
"gender": "male",
"age": "young",
"accent": "vietnamese",
"locale": "vi-VN",
"language": "vi",
"languages": ["vi"],
"use_case": "conversational",
"category": "high_quality",
"featured": true,
"usage_count": 12840
}
]
GET /samples/{name}/audio — Nghe thử giọng
Stream clip mẫu của một giọng để nghe trước khi tạo.
cURL
curl -sS "https://voice.saydi.ai/api/samples/vi-adam/audio" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
--output preview.wav
GET /v1/models — Liệt kê model
Danh sách model tương thích OpenAI: tts-1 (16 bước) và tts-1-hd (32 bước).
cURL
curl -sS "https://voice.saydi.ai/api/v1/models" \
-H "Authorization: Bearer $SAYDI_API_KEY"
POST /v1/voices
Đăng ký giọng một lần, dùng lại bằng voice_id.
Gửi lên một clip kèm tên. SaydiVoice tự kiểm tra audio, chọn đoạn nói liền mạch tốt nhất, phiên âm, và tách nhạc nền nếu phát hiện — rồi trả về voice_id ổn định. Dùng id đó trong POST /v1/audio/speech hoặc POST /tts: không phải upload lại clip, và audio mẫu được cache ở server nên các lần sinh sau nhanh hơn. Giọng đăng ký ở đây cũng hiện trong Studio → Giọng của tôi, và giọng tạo trong Studio cũng hiện ở đây.
POST /v1/voices — Đăng ký giọng
Upload multipart. Trả về object giọng nói; status là ready khi dùng được ngay.
| Name | Type | Required | Description |
|---|
file | file | yes | Clip mẫu: WAV, MP3, M4A, OGG, FLAC hoặc WebM. Dài tối đa 5 phút; hệ thống tự chọn đoạn tốt nhất. |
name | string | yes | Tên hiển thị. Không cần duy nhất — hãy định danh giọng bằng id, đừng dùng tên. |
consent | string | yes | Xác nhận bạn có quyền nhân bản giọng này. Gửi true, hoặc JSON {"accepted": true}. Được lưu kèm IP và phiên bản điều khoản. |
language | string | no | Gợi ý ISO-639-1 (ví dụ vi). Cũng quyết định ngôn ngữ của câu nghe thử. |
ref_text | string | no | Transcript của clip. Bỏ trống thì hệ thống tự phiên âm; gửi kèm (khi clip đã được chuẩn bị sẵn) để bỏ qua bước đó. |
separate | string | no | auto (mặc định) chỉ tách nhạc nền khi phát hiện; true luôn tách; false không tách. |
Idempotency-Key | string | no | Header. Gọi lại với cùng giá trị sẽ trả về giọng đã tạo thay vì tạo thêm — nên timeout không làm mất slot. |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/v1/voices" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "file=@my-recording.m4a" \
-F "name=Giọng của tôi" \
-F "language=vi" \
-F "consent=true"
# -> {"id":"vc_9f2c...","status":"ready","ref_text":"...","quality":{...}}
# Then speak with it — no clip to upload:
curl -sS -X POST "https://voice.saydi.ai/api/v1/audio/speech" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"vc_9f2c...","input":"Xin chào từ giọng của tôi."}' \
--output hello.mp3
Python
import os, requests
key = os.environ["SAYDI_API_KEY"]
base = "https://voice.saydi.ai/api"
with open("my-recording.m4a", "rb") as fh:
voice = requests.post(
f"{base}/v1/voices",
headers={"Authorization": f"Bearer {key}"},
files={"file": ("my-recording.m4a", fh, "audio/mp4")},
data={"name": "Giọng của tôi", "language": "vi", "consent": "true"},
timeout=300,
)
voice.raise_for_status()
voice = voice.json()
print(voice["id"], voice["status"], voice.get("quality", {}).get("warnings"))
audio = requests.post(
f"{base}/v1/audio/speech",
headers={"Authorization": f"Bearer {key}"},
json={"model": "tts-1", "voice": voice["id"], "input": "Xin chào từ giọng của tôi."},
timeout=180,
)
audio.raise_for_status()
open("hello.mp3", "wb").write(audio.content)
201 — Đã tạo giọng.
{
"id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12",
"object": "voice",
"name": "Giọng của tôi",
"status": "ready",
"language": "vi",
"ref_seconds": 7.4,
"ref_text": "Xin chào, đây là đoạn nói trong clip mẫu.",
"quality": {
"speech_seconds": 6.9,
"snr_db": 26.4,
"trimmed_seconds": 42.1,
"music_detected": false,
"separated": false,
"warnings": []
},
"consent": {"terms_version": "voice-terms-2026-01"},
"preview_url": "/v1/voices/vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12/preview"
}
200 — Gọi lại idempotent: giọng đã tồn tại và được trả về nguyên trạng.
{
"id": "vc_9f2c1a4b7e0d4a1c8b3f5e6d7a9c0b12",
"status": "ready",
"idempotent_replay": true
}
processing và ready — ready nghĩa là dùng được ngay — đó là trường hợp thường gặp với clip ngắn, sạch. processing nghĩa là clip có nhạc nền nên hệ thống còn đang tách: hãy gọi GET /v1/voices/{voice_id} (thường vài giây) trước khi sinh audio. Giọng failed sẽ trả lại slot và kèm error.code để xử lý. Hai lỗi nên xử lý riêng: 503 storage_unavailable là tạm thời — retry có backoff; 410 ref_missing nghĩa là audio đã mất hẳn, hãy đăng ký lại giọng (không tốn thêm slot).
GET /v1/voices — Liệt kê và quản lý giọng
Liệt kê giọng của bạn, lấy một giọng, đổi tên, xoá, hoặc nghe thử.
| Endpoint | Công dụng |
|---|
GET /v1/voices | This account's voices, newest first, with usage (slots used/limit). Filter with ?status=ready. |
GET /v1/voices/{voice_id} | One voice: status, ref_text and the quality report. |
PATCH /v1/voices/{voice_id} | Rename: JSON {"name": "..."}. The id never changes. |
DELETE /v1/voices/{voice_id} | Delete the voice and its stored audio, freeing its clone slot. Immediate. |
GET /v1/voices/{voice_id}/preview | MP3 of a short sentence in that voice. Rendered once and cached. Not charged against your characters. |
cURL
curl -sS "https://voice.saydi.ai/api/v1/voices" \
-H "Authorization: Bearer $SAYDI_API_KEY"
curl -sS -X PATCH "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Giọng nam trầm"}'
curl -sS "https://voice.saydi.ai/api/v1/voices/$VOICE_ID/preview" \
-H "Authorization: Bearer $SAYDI_API_KEY" --output preview.mp3
curl -sS -X DELETE "https://voice.saydi.ai/api/v1/voices/$VOICE_ID" \
-H "Authorization: Bearer $SAYDI_API_KEY"
Quyền sử dụng và hạn mức slot — Chỉ đăng ký giọng bạn có quyền sử dụng; API từ chối request thiếu consent. Mỗi gói có số slot nhân bản (clone_slots) — vượt hạn mức trả về 402 với code: voice_limit_reached. Đăng ký thất bại không tiêu tốn slot.
POST /tts/clone
Nhân bản và đọc trong một request.
Nhân bản và đọc trong một request multipart. Audio sạch, một người nói, dài 10–15 giây cho kết quả tốt nhất.
| Name | Type | Required | Description |
|---|
text | string | yes | Văn bản cần đọc bằng giọng nhân bản. |
ref_audio | file | yes | Clip mẫu. Nên là WAV mono; 10–15 giây audio sạch một người nói là tốt nhất. Tối đa 10 MB. |
ref_text | string | no | Transcript của clip. Rất nên có; tự phiên âm nếu bỏ trống. |
lang | string | no | Gợi ý ngôn ngữ cho chuẩn hoá và phiên âm. |
duration | number | no | Ép độ dài output, giây. Ghi đè speed. |
speed | number | no | Hệ số tốc độ. |
guidance_scale | number | no | Thang classifier-free guidance. |
num_step | integer | no | Số bước diffusion. |
breaks | string | no | Chuỗi JSON cấu hình auto-pause, ví dụ {"sentence":450,"comma":250}. |
output_format | string | no | wav, mp3, flac hoặc ogg. |
Python
import os, requests
with open("reference.wav", "rb") as fh:
r = requests.post(
"https://voice.saydi.ai/api/tts/clone",
headers={"Authorization": f"Bearer {os.environ['SAYDI_API_KEY']}"},
data={
"text": "Xin chào, đây là giọng đã nhân bản.",
"ref_text": "Đây là câu nói trong clip mẫu.",
"output_format": "mp3",
},
files={"ref_audio": ("reference.wav", fh, "audio/wav")},
timeout=180,
)
r.raise_for_status()
open("cloned.mp3", "wb").write(r.content)
Lưu ý — Audio mẫu tối đa 10 MB. Chỉ nhân bản giọng bạn có quyền sử dụng; upload lại clip mỗi lần gọi sẽ chậm hơn dùng giọng đã lưu.
POST /voice/transcribe
Phiên âm clip mẫu để nhân bản chính xác hơn.
Phiên âm clip mẫu ngắn để lấy ref_text, giúp nhân bản chính xác hơn.
| Name | Type | Required | Description |
|---|
audio | file | yes | Clip mẫu. Nên WAV mono, tối đa 10 MB. |
lang | string | no | Gợi ý ngôn ngữ tuỳ chọn (ISO 639-1). |
cURL
curl -sS -X POST "https://voice.saydi.ai/api/voice/transcribe" \
-H "Authorization: Bearer $SAYDI_API_KEY" \
-F "audio=@reference.wav" \
-F "lang=vi"
200 — Transcript kèm đánh giá chất lượng. Kiểm tra ok, không chỉ transcript.
{
"ok": true,
"reason": "ok",
"transcript": "Đây là câu nói trong clip mẫu.",
"word_count": 8,
"char_count": 25
}
Credits & quota
Cách tính phí text và nhân bản.
Text to speech và nhân bản tính theo ký tự, dùng chung một quota cho studio và API.
| Sự kiện | Meter | Có trừ không |
|---|
| POST /v1/audio/speech | characters | input length |
| POST /tts, /tts/file | characters | text length |
| POST /tts/clone | characters | text length |
| POST /voice/transcribe | feature gate | No |
| GET /samples, /v1/models | none | No |
Current plan and usage preflight
# Current subscription (plan, cycle, period end)
curl -sS "https://voice.saydi.ai/api/billing/me" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
# Would this text fit? Cheap check before spending GPU time.
curl -sS "https://voice.saydi.ai/api/billing/quota/preflight?chars=1200" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
# Everything the developer console shows, in one call (see Usage API below):
curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
API sử dụng
Truy vấn số dư, mức sử dụng và lịch sử gọi.
Cần access token của người dùng. Dùng các endpoint này cho dashboard và cảnh báo của bạn.
GET /usage/summary — Credits và số request
Số dư credits (dùng chung với studio) kèm số request API và chuỗi theo ngày.
| Name | Type | Required | Description |
|---|
days | integer | no | Khoảng thời gian của chuỗi, 1–365. Mặc định 30. |
cURL
curl -sS "https://voice.saydi.ai/api/usage/summary?days=30" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Credits lấy từ quỹ chung; requests/characters chỉ đếm traffic qua API.
{
"ok": true,
"plan_code": "pro",
"period_start": 1758844800.0,
"period_end": 1761436800.0,
"credits": {
"status": "ready",
"unlimited": false,
"limit": 600000,
"used": 41250,
"remaining": 558750,
"extra": 0,
"day_limit": null,
"day_used": 3200,
"day_remaining": null
},
"api": {
"limits": {"api_concurrency": 5, "clone_slots": 50, "max_chars_request": 100000},
"requests": {"total": 412, "success": 401, "errors": 11},
"characters": 41250,
"last_call_at": 1760001234.5,
"status": "ready",
"days": [
{"day": "2026-09-25", "requests": 18, "characters": 2100, "errors": 0},
{"day": "2026-09-26", "requests": 24, "characters": 3200, "errors": 2}
]
}
}
GET /usage/calls — Lịch sử gọi API
Các lần gọi mới nhất trước, kể cả lần bị từ chối, kèm cursor để phân trang.
| Name | Type | Required | Description |
|---|
limit | integer | no | Số bản ghi mỗi trang, 1–200. Mặc định 50. |
before | number | no | Cursor: giá trị next_before của trang trước. |
cURL
curl -sS "https://voice.saydi.ai/api/usage/calls?limit=25" \
-H "Authorization: Bearer $SAYDI_ACCESS_TOKEN"
200 — Dữ liệu lưu 400 ngày. next_before là null ở trang cuối.
{
"ok": true,
"status": "ready",
"next_before": 1760001000.0,
"calls": [
{
"ts": 1760001234.5,
"method": "POST",
"endpoint": "/v1/audio/speech",
"status": 200,
"chars": 41,
"voice": "vi-adam",
"latency_ms": 3820,
"key_id": "key_1f0c9a2b7d3e4f56",
"limit_reason": ""
},
{
"ts": 1760001000.0,
"method": "POST",
"endpoint": "/tts",
"status": 429,
"chars": 0,
"voice": "",
"latency_ms": 6,
"key_id": "key_1f0c9a2b7d3e4f56",
"limit_reason": "api-key-window"
}
]
}
Rate limit & đồng thời
Giới hạn request theo key và concurrency.
Giới hạn áp theo từng API key; concurrency theo gói.
| Giới hạn | Mặc định | Phản hồi khi vượt |
|---|
| Requests per key | 60 / minute | 429, Retry-After, X-OV-Limit: api-key-window |
| Concurrent generations | Pro 5 · Business 10 (trial 2) | 429, Retry-After, X-OV-Limit: api-concurrency |
Retry with backoff
import time, requests
def speak(text, voice, tries=4):
for attempt in range(tries):
r = requests.post(
"https://voice.saydi.ai/api/v1/audio/speech",
headers={"Authorization": f"Bearer {KEY}"},
json={"input": text, "voice": voice},
timeout=180,
)
if r.status_code != 429:
r.raise_for_status()
return r.content
wait = int(r.headers.get("Retry-After", "5"))
time.sleep(max(1, wait) * (2 ** attempt)) # exponential backoff
raise RuntimeError("rate limited after retries")
Khi gặp 429 — Chờ theo Retry-After, rồi retry với backoff và jitter.
Mã lỗi
Dạng lỗi, mã thường gặp và cách retry.
Endpoint native trả {"detail": ...}. POST /v1/audio/speech và GET /v1/models trả dạng OpenAI {"error": {...}} cho MỌI lỗi — sai credential (401), field sai (422), hết hạn mức hoặc quá nhanh (429), lỗi server (500) — và giữ nguyên các header của response (Retry-After, X-OV-Limit). Một khác biệt có chủ ý: field sai ở đây giữ status 422, còn OpenAI trả 400 cho cùng lỗi đó.
Native
{
"detail": "Field 'text' must not be empty."
}
OpenAI-compatible
{
"error": {
"message": "Field 'input' must not be empty.",
"type": "invalid_request_error",
"param": "input",
"code": "invalid_value"
}
}
| Status | Ý nghĩa | Thử lại? |
|---|
| 400 | Malformed request, empty text, or undecodable reference audio | No — fix the request |
| 401 | Missing, unknown, revoked or expired credential | No — fix the credential |
| 403 | Account or token blocked for suspected abuse | No — contact support |
| 404 | Endpoint or voice not found | No |
| 413 | Text too long, or upload above 10 MB | No — split or compress |
| 422 | Request failed schema validation (wrong type or range) | No — check parameter types |
| 429 | Quota exhausted, per-key rate limit, or concurrency cap (on /v1/*, error.code separates them: insufficient_quota vs rate_limit_exceeded) | Yes — after Retry-After |
| 500 | Unexpected server error | Yes — with backoff |
| 503 | Backend temporarily unavailable or not configured | Yes — with backoff |
X-OV-Limit trên 403/429 cho biết limit nào đã chặn request.
| X-OV-Limit | Status | Nguyên nhân |
|---|
| api-key-window | 429 | Per-key request rate exceeded |
| api-concurrency | 429 | Too many in-flight generations for the plan |
| usage-limit | 429 / 403 | Credit quota exhausted, or account soft-limited |
| blocked | 403 | Account or token blocked |
| ip-window | 429 | Edge per-IP limit (shared infrastructure, not your key) |
| concurrency | 503 | Server-wide in-flight ceiling reached |
Changelog
Thay đổi gần đây và chính sách phiên bản.
| Ngày | Phiên bản | Thay đổi | Breaking? |
|---|
| 2026-09-26 | 1.0 | Giai đoạn 1: per-account API keys, per-key rate limit, per-plan concurrency, shared character credits. | — |
Chính sách phiên bản — Thay đổi cộng thêm không cần báo trước. Thay đổi breaking sẽ có prefix /vN mới và ghi tại đây.
Hỗ trợ
Kênh hỗ trợ.
| Chủ đề | Kênh |
|---|
| Integration questions | contact@saydi.ai |
| Higher limits / Enterprise | contact@saydi.ai |
| Suspected abuse or a leaked key | Revoke the key in the dashboard, then contact support |
| Service status | GET /health (no auth required) |