# Add file-based voice input and spoken replies to your OpenRouter app

> Build file-based voice input and replies with Scribe v2 and Eleven v4 Turbo; measure both costs, latency, and transcript quality.
> By Dave · 2026-10-08
> Source: https://otf-kit.dev/blog/elevenlabs-on-openrouter-voice-app

A voice feature is not one API call. It is a sequence of audio input, transcription, application logic, and spoken output, each with its own latency, format, and cost behavior. OpenRouter’s October 7 launch makes ElevenLabs speech-to-text and text-to-speech models callable alongside an app’s existing OpenRouter model traffic. That can simplify routing and billing, but it does not turn endpoint-based audio into a real-time voice connection.

For an app that already uses OpenRouter, the useful first slice is narrow: transcribe a recorded user utterance with Scribe v2, pass the text to the app’s existing chat model, and speak the response with Eleven v4 Turbo. Test that path on audio you own before putting it behind a microphone button.

![Dex and Nova compare an audio upload, character-count billing, and a playable MP3 response](https://cdn.otf-kit.dev/blog/elevenlabs-on-openrouter-voice-app/inbody1-audio-cost-and-format-20261008a.png)

## What the launch adds

OpenRouter lists nine ElevenLabs text-to-speech models and two speech-to-text models. The launch’s main building blocks are `elevenlabs/scribe-v2` for transcription, `elevenlabs/eleven-v4-turbo` for lower-latency spoken replies, and `elevenlabs/eleven-v4` for more expressive narration. The announcement says v4 Turbo is intended for real-time use and bills at half the v4 per-character rate. That describes the model’s relative price tier; measure your actual usage and current rate before estimating a feature’s bill.

Scribe v2 returns a transcript and can provide word timestamps, speaker labels, and audio-event tags. Its request can report audio seconds and a dollar cost in a `usage` object. Text-to-speech is billed by input characters, counted as Unicode code points; audio tags count too. These are different meters. A longer recording increases transcription duration, while a longer generated script increases speech input characters.

The distinction matters for product design. A short command can still come from a long recording with pauses. A short recording can lead to a long spoken answer. Track both sides separately instead of treating “voice usage” as one number.

## Start with a recorded utterance

First, send a short audio sample to Scribe v2 and inspect the transcript. OpenRouter’s launch walkthrough shows JSON input with base64-encoded audio and an explicit source format. For a file-based workflow, it also shows multipart upload with `file`, `model`, and response-format fields. The examples are alternatives; select the request shape that matches how your client stores or uploads audio.

A minimal multipart call for a local WAV file looks like this:

```bash
curl https://openrouter.ai/api/v1/audio/transcriptions \
  --fail-with-body \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F file="@question.wav" \
  -F model="elevenlabs/scribe-v2"
```

Keep the key on your server. The client should upload audio to an authenticated application endpoint, and that server should call OpenRouter. Enforce an upload-size and duration limit before forwarding the file; the launch says uploaded audio is capped at 25 MB. Its stated cap covers about 27 minutes of 128 kbps MP3, but the request can still time out upstream after 180 seconds. A product should set a smaller practical maximum based on its user flow and measured latency.

For a first release, avoid streaming claims. This launch does not include realtime Scribe v2 over WebSocket. If the interaction needs continuous back-and-forth audio, compare it with the separate realtime path described in [our guide to Gemini Live on AI Gateway](/blog/gemini-3-8-live-ai-gateway). The two workflows solve different product problems: one processes request/response audio files, while the other is a realtime session.

## Make the spoken response playable

After transcription, pass the text to the chat model already used by your app. Then send the answer to the speech endpoint. Set `response_format` to `mp3` when you need a file most browser and mobile players can open:

```bash
curl https://openrouter.ai/api/v1/audio/speech \
  --fail-with-body \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs/eleven-v4-turbo",
    "input": "Your appointment is confirmed for Tuesday at two.",
    "voice": "sarah",
    "response_format": "mp3"
  }' \
  --output reply.mp3
```

The announcement says that omitting `response_format` returns raw 16-bit little-endian mono PCM at 24 kHz. That may suit a custom audio pipeline, but it is not the default to choose for a downloadable voice reply. With `mp3`, the announcement describes 44.1 kHz, 128 kbps output. Verify the response content type and playback on each client you support; a successful HTTP response alone does not prove the output can be played by the interface.

Eleven v4 Turbo has no speed control and rejects a non-default `speed` setting. For that model, adjust phrasing or use supported delivery tags where appropriate. Tags are text sent to the model and count toward the character meter. Don’t apply options copied from another ElevenLabs model without checking the selected model’s supported fields.

## Measure the two cost meters

For transcription, record audio duration and the response’s reported `usage.cost` for every test request. Summarize cost by duration bands that match your product (for example, short commands and longer notes), not only as a blended average. Include failed, retried, and abandoned requests in your own product metrics so provider usage is not hidden by a successful-response dashboard.

For speech output, record the number of Unicode code points sent, including tags, and the selected model. The per-character rate differs by model tier; the launch says v4 Turbo is half the v4 rate. A temporary launch discount is also described on the announcement, with an end time of October 19, 2026 at 8 a.m. Pacific Time. Treat that as a dated promotion, not a standing price assumption. Recheck the model’s live pricing when setting limits or projecting monthly spend.

A useful per-feature estimate is requests per active user multiplied by measured transcription cost and measured speech cost, then separated by model and duration or character band. Keep the raw counts in your analytics so the estimate can be recomputed when a rate changes. Set an application-level budget and a maximum response length before wider rollout.

![Byte and Luna review a sample recording with speaker-colored transcript tracks and word timestamps](https://cdn.otf-kit.dev/blog/elevenlabs-on-openrouter-voice-app/inbody2-transcript-quality-speaker-check-20261008a.png)

## Test quality and speaker labels on your own samples

A transcript that reads well in a clean demo may fail on the audio your users actually send. Build a small evaluation set from consented or otherwise approved recordings that represent your expected microphones, accents, room noise, speaking pace, and vocabulary. Keep a reference transcript for each clip. Compare output against it, review proper nouns and numbers separately, and listen to the generated reply rather than checking only the text response.

For meeting notes or multi-person recordings, request `verbose_json` and word-level timestamps, then enable diarization in the provider options. The launch example also passes a known speaker count and audio-event tagging. Each word entry can include the recognized word, start and end times, confidence, speaker number, and speaker label. Treat those labels as model output to review: a `speaker_0` label is not a verified identity. If your application must name people, map labels only through a separate, explicit user-confirmed step.

Test clips with speakers interrupting, overlapping, or sitting at different distances from the microphone. Track transcription errors that change intent, not just average word error. A wrong “can” versus “can’t,” amount, date, or product name can matter more than several harmless filler-word differences. For generated audio, check that the file starts, ends, and plays on supported devices, and measure time to first playable audio as well as total response time.

## Measure latency before calling it conversational

These endpoints form a serial path: upload, transcription, chat response, speech generation, and audio delivery. Log a timestamp at each boundary so a slow interaction has an identifiable cause. Measure median and tail latency on the same network conditions and with the same sample set. Include audio duration, transcript length, generated character count, chosen model, and whether the request was retried.

Set a product threshold for when to show progress, when to allow cancellation, and when to offer a text alternative. If users need the system to listen while they speak and answer with minimal turn-taking delay, the endpoint chain may not fit. Do not infer realtime behavior from the phrase “built for real-time use” on v4 Turbo: the OpenRouter example still makes a request and receives a generated file, and realtime Scribe over WebSocket is explicitly outside this launch.

## A cautious rollout

Start with a server-side prototype that accepts one short recording, transcribes it, sends the text through your existing model call, and returns an MP3. Use fixed samples first; then invite a small group to try it with clear recording and retention expectations. Review quality, latency, and both cost meters before broadening the duration limit or enabling diarization by default.

OpenRouter says these calls use the same API key as other OpenRouter model requests, and the announcement says a separate ElevenLabs plan is not required for the hosted models. If your product depends on cloned voices, pronunciation dictionaries, or higher-quality output formats, the announcement points existing ElevenLabs users to its bring-your-own-key path. Confirm the account, privacy, and retention terms your product needs before routing user audio through any hosted provider.

The practical decision is straightforward: use Scribe v2 and v4 Turbo for a measured, file-based voice feature when their request/response behavior fits your interface. Keep the transcription seconds, generated characters, latency, and sample quality visible in your own metrics. Choose a realtime architecture when the product requires an ongoing audio session, rather than stretching these endpoints into one.

## Sources

- [OpenRouter: ElevenLabs is now on OpenRouter](https://openrouter.ai/blog/announcements/elevenlabs-on-openrouter/) — model choices, request examples, request limits, output format, billing units, and the launch promotion.
