OpenRouter adds unified audio transcription

Canonical version: OpenRouter adds unified audio transcription.

OpenRouter added speech-to-text to its API. A new endpoint (POST /api/v1/audio/transcriptions) accepts base64-encoded audio and returns the transcript as JSON, using the same API key as chat completions, so you don't need a separate STT provider or server.

Two model families are available: Whisper-class models priced per audio second, and newer STT models priced per token (discoverable with the ?output_modalities=transcription filter). Pricing is the model's catalog rate with no OpenRouter markup, and every response includes a usage object with duration, token counts, and the actual cost of that request. When a model is hosted by several providers, requests load-balance automatically.

Current limits: a 60-second processing timeout (a throughput constraint; audio length itself isn't capped), no URL-based audio input (you upload the bytes), and no built-in SRT/VTT subtitle output.

For anyone already routing Large Language Models (LLMs) traffic through OpenRouter, this makes meeting transcripts, voice notes, and podcast-to-text pipelines a one-endpoint addition instead of a second vendor relationship.

References


About Sébastien

Ready to get to the next level?

Found this valuable? Share it with someone who needs it.

Join 6,000+ readers. Get practical systems for knowledge & AI. Free.

Subscribe ✨

Free: Knowledge System Checklist

A clear roadmap to building your own knowledge system. Subscribe and get it straight to your inbox.

6,000+ readers. No spam. Unsubscribe anytime.

Subscribe