Skip to main content
FreeWhat would your company’s AI control layer look like?Build mine

Voice Gateway

Speech in. Speech out.
Same gateway.

Run ElevenLabs text to speech and speech to text through Hicap. Same base URL, same api-key header, same bill as the rest of your AI traffic.

Model IDs on the gateway

eleven_v3eleven_multilingual_v2scribe_v1scribe_v2

POST /v1/text-to-speech·eleven_v3

Text to speech

“The first move is what sets everything in motion.”

00:02.4 / 00:04.1audio/mpeg

POST /v1/speech-to-text·scribe_v2

Speech to text

S100:00.4Can we move the launch to Thursday?

S200:02.1Thursday works. I will tell the team.

Both on api.hicap.ai/v1 · one api-key header · one bill

Integration

Nothing about your
integration changes.

Voice is consolidation, not a second setup track. Teams already on Hicap should not have to think about it as another platform.

chat
api.hicap.ai/v1
voice
api.hicap.ai/v1

Same base URL.

Voice runs on the endpoint your chat traffic already uses.

api-key: $HICAP_API_KEY

one credential, both directions

Same api-key header.

Authenticate with the Hicap key you already have.

POST /v1/text-to-speech/{voice_id}
"model_id": "eleven_v3"

ElevenLabs request shape.

Use the ElevenLabs voice paths and model IDs you already call.

Voice analytics

Tracked to the minute.
Next to everything else.

Spend, audio hours, requests, and cost per minute for both directions, in the same dashboard as chat. No second vendor to meter or reconcile.

  • STT and TTS, side by side

    Flip between transcription and generation and every readout re-scopes.

  • When it happens

    A day-by-hour heat map of requests or spend shows where the load really sits.

  • Cost per minute

    Daily cost and average cost per audio minute, the unit voice is actually priced in.

Voice Analytics

STTTTS

$56.70

STT spend

172.35

Audio hours

1.0K

Requests

620.5s

Avg / request

Activity Heat Map

Day of week × hour of day (UTC)

RequestsSpend
Sun
Mon
Tue
Wed
Thu
Fri
Sat
036912151821

STT Daily Cost

Per-day spend in window

Apr 12Apr 22May 2May 12May 22Jun 1Jun 11

Models

Choose the right voice model.

Expressive voice work, steady long-form narration, and production transcription, all from one route.

ModelModeBest forHighlights

Eleven v3

eleven_v3

TTSExpressive, character-driven voice work70+ languages, multi-speaker dialogue
Up to 5,000 characters per request

Eleven Multilingual v2

eleven_multilingual_v2

TTSSteady long-form narration and explainers29 languages, consistent delivery
Up to 10,000 characters per request

Scribe v1

scribe_v1

STTSearchable text from audio and video90+ languages
Word-level timestamps

Scribe v2

scribe_v2

STTProduction transcripts that need accuracyDiarization up to 32 speakers, audio tagging
Keyterm prompting up to 1,000 terms, optional cleanup

Quickstart

Same URL.
Same api-key header.

Keep the ElevenLabs endpoint shapes and model IDs. Authentication and routing move onto Hicap.

1curl --request POST \
2 --url "https://api.hicap.ai/v1/text-to-speech/JBFqnCBsd6RMkjVDRZzb" \
3 --header "Content-Type: application/json" \
4 --header "api-key: $HICAP_API_KEY" \
5 --data '{
6 "text": "The first move is what sets everything in motion.",
7 "model_id": "eleven_v3"
8 }' \
9 --output speech.mp3

Replace JBFqnCBsd6RMkjVDRZzb with the ElevenLabs voice ID you want. Other response formats follow the ElevenLabs request options.

Voice coverage today: Eleven v3 and Eleven Multilingual v2 for generation, Scribe v1 and Scribe v2 for transcription. The rest of the model catalog stays on the same account.