Skip to main content
Use the audio endpoints when the result is a file or text. Use the realtime WebSocket endpoint for a live, two-way audio experience.

Choose an endpoint

Choose a model

Read the current model list instead of hard-coding it. Check the selected model’s details for realtime support before opening a socket.

Audio requests

Speech returns audio in the HTTP response. Transcription and translation can return final text or an accepted task. When the body contains a task ID and status instead of final text, save the ID and follow poll_url until completion or failure. Allow enough time for large files and keep the Request ID. Eligible generated audio HTTP(S) result URLs may be retained as media copies for 30 days. Check media_retention.items for each item’s status and expires_at; pending or failed copies are not guaranteed. Inline or binary audio responses are outside this URL-copy retention. See Data retention.

Realtime Sessions

Open a WebSocket with the model in the query string and the API key in the Authorization header. Keep the event format documented for the selected realtime model, and close the socket when the session is complete. This guide covers the WebSocket subset only. TokenLab does not currently provide OpenAI Realtime REST client secret, translation client secret, Calls, or legacy beta session management endpoints. For the server-side example below, install ws and set TOKENLAB_REALTIME_MODEL to a current model whose detail declares /v1/realtime. Provide TOKENLAB_API_KEY in the server environment; the name alone does not establish realtime support.

Show progress safely

  • Save generated audio files instead of replaying the same request on refresh.
  • For transcription and translation, show upload and processing states even when the API call is synchronous.
  • For realtime, handle close events and reconnect only after the user starts a new session.
  • Do not put API keys, private URLs, or account secrets in audio text input.

API Reference