OpenRouter adds unified audio transcription
Canonical version: OpenRouter adds unified audio transcription.
OpenRouter added speech-to-text to its API. A new endpoint (POST /api/v1/audio/transcriptions) accepts base64-encoded audio and returns the transcript as JSON, using the same API key as chat completions, so you don't need a separate STT provider or server.
Two model families are available: Whisper-class models priced per audio second, and newer STT models priced per token (discoverable with the ?output_modalities=transcription filter). Pricing is the model's catalog rate with no OpenRouter markup, and every response includes a usage object with duration, token counts, and the actual cost of that request. When a model is hosted by several providers, requests load-balance automatically.
Current limits: a 60-second processing timeout (a throughput constraint; audio length itself isn't capped), no URL-based audio input (you upload the bytes), and no built-in SRT/VTT subtitle output.
For anyone already routing Large Language Models (LLMs) traffic through OpenRouter, this makes meeting transcripts, voice notes, and podcast-to-text pipelines a one-endpoint addition instead of a second vendor relationship.
References
Related
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.