Voice Gateway
Speech in. Speech out.
Same gateway.
Run ElevenLabs text to speech and speech to text through Hicap. Same base URL, same api-key header, same bill as the rest of your AI traffic.
Model IDs on the gateway
POST /v1/text-to-speech·eleven_v3
Text to speech
“The first move is what sets everything in motion.”
POST /v1/speech-to-text·scribe_v2
Speech to text
S100:00.4Can we move the launch to Thursday?
S200:02.1Thursday works. I will tell the team.
Both on api.hicap.ai/v1 · one api-key header · one bill
Integration
Nothing about your
integration changes.
Voice is consolidation, not a second setup track. Teams already on Hicap should not have to think about it as another platform.
Same base URL.
Voice runs on the endpoint your chat traffic already uses.
one credential, both directions
Same api-key header.
Authenticate with the Hicap key you already have.
ElevenLabs request shape.
Use the ElevenLabs voice paths and model IDs you already call.
Voice analytics
Tracked to the minute.
Next to everything else.
Spend, audio hours, requests, and cost per minute for both directions, in the same dashboard as chat. No second vendor to meter or reconcile.
STT and TTS, side by side
Flip between transcription and generation and every readout re-scopes.
When it happens
A day-by-hour heat map of requests or spend shows where the load really sits.
Cost per minute
Daily cost and average cost per audio minute, the unit voice is actually priced in.
Voice Analytics
$56.70
STT spend
172.35
Audio hours
1.0K
Requests
620.5s
Avg / request
Activity Heat Map
Day of week × hour of day (UTC)
STT Daily Cost
Per-day spend in window
Models
Choose the right voice model.
Expressive voice work, steady long-form narration, and production transcription, all from one route.
| Model | Mode | Best for | Highlights |
|---|---|---|---|
Eleven v3 eleven_v3 | TTS | Expressive, character-driven voice work | 70+ languages, multi-speaker dialogue Up to 5,000 characters per request |
Eleven Multilingual v2 eleven_multilingual_v2 | TTS | Steady long-form narration and explainers | 29 languages, consistent delivery Up to 10,000 characters per request |
Scribe v1 scribe_v1 | STT | Searchable text from audio and video | 90+ languages Word-level timestamps |
Scribe v2 scribe_v2 | STT | Production transcripts that need accuracy | Diarization up to 32 speakers, audio tagging Keyterm prompting up to 1,000 terms, optional cleanup |
Quickstart
Same URL.
Same api-key header.
Keep the ElevenLabs endpoint shapes and model IDs. Authentication and routing move onto Hicap.
1curl --request POST \2 --url "https://api.hicap.ai/v1/text-to-speech/JBFqnCBsd6RMkjVDRZzb" \3 --header "Content-Type: application/json" \4 --header "api-key: $HICAP_API_KEY" \5 --data '{6 "text": "The first move is what sets everything in motion.",7 "model_id": "eleven_v3"8 }' \9 --output speech.mp3
Replace JBFqnCBsd6RMkjVDRZzb with the ElevenLabs voice ID you want. Other response formats follow the ElevenLabs request options.
Voice coverage today: Eleven v3 and Eleven Multilingual v2 for generation, Scribe v1 and Scribe v2 for transcription. The rest of the model catalog stays on the same account.