Quincer AI Docs Home Support Start free

Configure

Voice chat

Visitors can switch to a spoken conversation with your agent at any point. A single tap on the headphone chip starts a real-time voice call through xAI’s Grok voices. Interruption, transcript capture, and graceful fallbacks to text are all handled for you.

Turn it on

Almost all of setup lives on one screen: Settings → Voice. In the figure below, the Call behavior card sets the Silence timeout and Max call length and picks a Default voice provider; the Voice API keys card underneath is where each provider either Reuses your text-chat key or takes a separate voice key — a green Ready badge appears the moment a usable key is on file. Only the per-widget Enable voice chat toggle and layout live elsewhere (on the Customize page — see below).

Settings · Voice
×
Call behavior

Defaults applied to every voice session — changes take effect on the next call.

Silence timeout (seconds)
10
Range 3–60s
Max call length (seconds)
900
Default 15 min
Default voice provider
xAI (Grok)
Voice API keys
xAI (Grok Voice) Ready
Reuse my text-chat XAI keyEnter a separate voice key

Voice keys are encrypted at rest. A Ready badge means a usable key is on file — your own, or the platform-default xAI key.

The Settings → Voice tab — call behavior on top, per-provider keys below; Ready confirms a usable key.

  1. Open Settings → Voice in the dashboard.
  2. Pick your provider (xAI Grok, OpenAI Realtime, or Google Gemini Live) and how to authenticate. The simplest option is Reuse my text-chat key — if text chat already works with that provider, voice just works too. Otherwise paste a voice-enabled key (console.x.ai for xAI, platform.openai.com for OpenAI, aistudio.google.com for Gemini).
  3. Pick a silence timeout (default 10s) and max call length (default 15 min).
  4. Open Billing → Voice and pick a plan (see below). Metered customers get minute bundles; BYOK flat gets unlimited minutes for a flat $49/mo platform fee.
  5. In Customize & Deploy → Appearance, pick a voice UI layout and flip the Enable voice chat switch for each widget you want the chip to appear on. Voice stays hidden on widgets where the toggle is off.
  6. Click Try voice live on the same page to test the selected layout and voice end-to-end in an embedded preview before visitors ever see it.
i

You don’t need your own xAI key to try voice — Quincer AI ships a platform-default key that works out of the box on every plan. Pasting your own only matters if you want to bill xAI directly (BYOK) or use a region-specific key.

Pricing

Two billing models, both requiring you to bring your own xAI API key:

Model Minutes Price Best for
Voice 60 Monthly 60/mo $26/mo Trying voice on a single widget
Voice 180 Monthly 180/mo $59/mo Moderate use across a couple of widgets
Voice 960 Monthly 960/mo $249/mo High-volume support / lead teams
Voice 60 / 180 / 960 Add-on one-time $29 / $69 / $299 Top up a busy period without changing your plan
Voice BYOK (unlimited) Unlimited $49/mo You pay xAI directly for usage; we charge a flat platform fee

Monthly plans are ~10–17% cheaper than the equivalent one-time add-on.

UI layouts

Voice is visually distinct from text, so Quincer AI offers four layouts for how the call is presented inside the widget. Both live on the Customize & Deploy page: the Enable voice chat toggle is on the AI agent tab (it needs a voice plan on the widget’s owner), and once it’s on, the Voice UI layout card appears with the four-way picker below. In the figure, the dropdown carries the exact labels the product shows — Hybrid (recommended), Takeover, Inline composer, and Compact strip — and a live preview under it renders the chosen layout before you save. The Try voice live button (top-right of the preview panel) opens the real widget so you can test the layout and voice end-to-end.

Customize & Deploy
×
Enable voice chat
Shows the voice chip in this widget’s header. AI agent tab · requires a voice plan.
Voice UI layout

How voice appears when a visitor activates it. Mobile widths auto-promote Hybrid to Takeover.

Hybrid (recommended)

Also: Takeover · Inline composer · Compact strip. A live preview renders the chosen layout below the dropdown.

Try voice liveSave

On Customize & Deploy — enable voice on the AI agent tab, then pick one of four layouts; Try voice live tests it end-to-end.

A live preview shows exactly how your chosen layout looks before you save.

LayoutWhat it looks likeBest for
Hybrid default A compact voice strip slides in above the chat transcript. Transcripts for both sides render inline in the message list as italic bubbles with a mic glyph, so visitors can scroll back through what was said. Most widgets. Works on desktop and mobile; auto-promotes to takeover under 480px.
Takeover The whole chat panel swaps to an orb-centric voice UI while the call is live. Text history is hidden until the call ends. Maximum focus. Support flows where voice is the primary channel, or mobile-first deployments.
Inline No dedicated strip or orb — transcripts simply appear in the normal message list while a mic-state dot in the header indicates listening / thinking / speaking. Minimal UIs that want voice to feel like dictation rather than a “call”.
Strip A slim always-visible control strip at the top of the chat when a call is active. Transcripts render in the message list like Inline, but with richer on-strip controls. Teams that want the call state obvious without taking over the whole panel.
i

On screens narrower than 480px, Hybrid and Strip automatically promote to Takeover — there’s not enough room to show a transcript and voice controls side-by-side. The dashboard preview has a Desktop/Mobile toggle so you can see both states.

How it works

Interrupting the opening prompt

When a call connects, your agent speaks an opening prompt. While that prompt is playing the caller’s microphone is ignored, so a door closing, a colleague talking nearby, or hold music bleeding in can’t cut the greeting off half way through. This is on by default and it is what we recommend.

It matters more than it sounds. Without it, the speech detector can trigger on the greeting itself or on background noise, and the call starts mid-sentence with the agent already listening to nothing. On one live call that misfire was transcribed as a different language and locked the whole conversation to it.

You’ll find the control in Customize → Deploy, directly under Enable voice chat:

i

You don’t need to know how long your prompt is. The setting can only make the mic open earlier — never later. If your prompt finishes before the time you set, it finishes first and your setting simply has no effect. Leave the seconds field blank to keep the mic off for the whole prompt.

Why a 3-second minimum? Below that, ordinary background noise interrupts the prompt before the caller has heard enough of it to know who they’ve reached — which is the problem the protection exists to solve.

Unticking the box without setting a number means barge-in from the very first millisecond. That has always been available and is unchanged; the seconds field is the middle ground between the two.

Live voice takeover

When an AI voice call needs a human, any licensed teammate can take it over mid-conversation. Visitor voice stays on Quincer AI’s relay while the AI is talking; the moment an agent clicks Take Over on the Live Conversations page, the widget bridges the visitor onto Cloudflare’s WebRTC SFU and the agent joins the same call. No app install for the visitor, no phone number, no third-party dial-in.

In the figure, the open conversation carries a Voice call badge in its header while the AI is live, and the action bar at the bottom shows the single Take Over button. If you don’t hold a live-voice seat, clicking it surfaces the amber gate notice shown here instead of the call console.

Live Conversations Voice call
×
Visitor · live voice
AI is on the call — grounded in your knowledge base
This visitor is on a live voice call. You don’t have a live voice seat — ask them to switch to text, or have an admin assign you a seat in Team settings.
Take Over

A live voice call on Live Conversations — the Voice call badge, one Take Over button, and the seat-gate notice.

  1. On Live Conversations the active voice call shows a Voice call badge. Open the conversation.
  2. Click Take Over. If you have a live-voice seat (see below), you’re redirected to the call console; if not, you get a banner suggesting the visitor be asked to switch to text.
  3. The AI speaks a short hand-off line (“Stand by — Mo is joining the call now”) while your console auto-requests mic access.
  4. Within a few seconds, your console and the visitor’s widget are bridged directly via Cloudflare Realtime. The AI steps aside. You and the visitor hear each other in real time.
  5. Either side clicking Hang up (or the visitor clicking End call) ends the bridge and closes the call for the other party.
i

If the WebRTC bridge fails (for example, firewalled networks), the AI stays on the call and tells the visitor: “I’m getting an error transferring you — let me stay on and help out, or switch to text and my colleague will chat with you there.” The session continues; no dropped call.

Seats & licensing

Live voice takeover is a per-seat SKU at $19/user/month. Assign seats on OrganizationPeople, using the Voice seat toggle on each teammate’s row. Super-admins can also grant seats directly from AdminUsers for free while you’re setting up. Visitor voice minutes continue to bill from the org’s voice plan; the $19 seat covers only the human handover side.

Keypad entry (touch-tones)

On phone calls, callers can answer with the keypad instead of their voice — an account number, an order number, a PIN. Your persona receives the digits as caller input, answers them like any spoken turn, and the entry is saved to the conversation transcript so you can see exactly what was keyed.

An entry is complete when any of these happens:

A caller who is mid-entry is a caller who is present: typing resets the silence timeout, so keying a long number never trips the “are you still there?” wind-down. And if the persona is mid-sentence when an entry lands, it stops talking and answers the entry.

Press 0 for a person

A lone 0 followed by that same ~2.5-second pause is not data — it’s the universal “get me a human”, and it triggers the same transfer-to-a-human flow as asking out loud: a warm transfer to the persona’s configured transfer target, or a team notification when none is set. One escalation per call — repeated 0 presses after the first are ignored.

A 0 inside a longer entry is just a zero. Account numbers and phone numbers contain zeros routinely; entering 5501# sends the persona 5501 — it never transfers the call.

+1 415 555 0142 Voice call
×
Agent
Sure — what are the last four digits of the account number?
Visitor
[The caller entered this on their phone keypad: 4821]

The caller pressed 4 8 2 1 ## submitted the entry immediately, and the transcript records exactly what the persona received.

i

On Google Gemini Live, keypad entries can’t be delivered to the model (its protocol has no channel for them) — callers can still speak the digits. Pressing a lone 0 to reach a person works on every provider: it’s handled before the AI ever sees it.

Voice providers

Quincer AI ships with three production voice providers — xAI Grok, OpenAI Realtime, and Google Gemini Live. Pick a default provider under Settings → Voice → Default voice provider; individual personas can select any voice from any enabled provider. All three run over the same relay pipeline, so everything else (layouts, billing, transcripts, knowledge grounding, integration tools) behaves identically.

ProviderModelGood for
xAI Grok grok-voice-think-fast-1.0 default · grok-voice-fast-1.0 Fast, conversational, five curated voices. Grok Think Fast is the flagship — background reasoning, stronger tool-calling, and 20+ languages including Arabic (Egypt, Saudi Arabia, UAE), Bengali, Hindi, Korean, Turkish, Vietnamese, and more. Legacy Grok Fast is still selectable in Settings → Voice while xAI completes its migration.
OpenAI Realtime gpt-realtime Ten voices including the expressive marin and cedar. Strongest when you need nuanced delivery.
Google Gemini Live gemini-3.1-flash-live Eight voices (Aoede, Charon, Fenrir, Kore, Leda, Orus, Puck, Zephyr). 70+ languages out of the box, strong multilingual delivery.

Pick the OpenAI realtime model

Running voice on OpenAI? The model is yours to choose. When your Default voice provider is OpenAI Realtime (gpt-realtime), an OpenAI realtime model field appears directly beneath the provider picker on Settings → Voice. Type any realtime model id — including a new release the day it ships. Leave it blank to stay on the platform default (gpt-realtime); the help text under the field always shows which model calls currently use.

Settings · Voice
×
Default voice provider
OpenAI Realtime (gpt-realtime)
OpenAI realtime model
gpt-realtime
Calls currently use gpt-realtime. Leave blank for the platform default (gpt-realtime). Must be a realtime speech-to-speech model — a model id OpenAI doesn’t recognise fails when the call connects.

The OpenAI realtime model field on Settings → Voice — it appears once the default provider is OpenAI Realtime; blank means the platform default.

Two things to know before you type a custom id. It must be a realtime speech-to-speech model — a model id OpenAI doesn’t recognise fails when the call connects, not when you save. And billing follows the model: voice cost is looked up by model id, so the model you pick is the model that answers and the one your usage reflects. Any custom id shows a Check pricing for this model reminder: if we hold no price for it yet, calls still work but their cost reports as $0 until a price is added. For the full model catalogue and how model pricing works across the product, see AI models.

Voices

Each persona picks one voice on its Voice & phone card, under Personas → [Persona]. In the figure, the Voice (optional) dropdown groups voices by provider (xAI, OpenAI Realtime, Gemini Live) and the first option is Provider default; the Preview button beside it plays a short sample so you can compare tone before saving. If you pick a voice from a provider that isn’t your workspace’s active one, an inline Mismatch warning appears with a link to Settings → Voice — Preview still works so you can audition any voice.

Voice & phone
×

The call voice, spoken phone greeting, and default call language for this persona.

Voice (optional)
Eve — warm, balanced female
Preview

Groups: xAI (Grok Voice) · OpenAI Realtime · Google Gemini Live. First option is Provider default (Eve / alloy / Aoede).

✓ Compatible with your XAI voice workspace

A persona’s Voice & phone card — a provider-grouped dropdown plus a Preview button to audition tone before saving.

xAI Grok

VoiceTone
Eve xAI defaultWarm, balanced female voice. Good all-rounder.
AraSofter female, calmer cadence.
LeoFriendly male, conversational.
RexDeeper male, more confident delivery.
SalNeutral, professional — works well for support.

OpenAI Realtime

VoiceTone
alloy OpenAI defaultBalanced, versatile. OpenAI’s neutral-tone pick.
marin recommendedMost natural delivery. Recommended for nuanced conversations.
cedar recommendedMost natural delivery. Slightly richer than marin.
ashExpressive, natural male.
balladWarm, storytelling.
coralFriendly, upbeat female.
echoClear, confident male.
sageCalm, thoughtful.
shimmerBright, energetic female.
verseSmooth, musical.

Google Gemini Live

VoiceTone
Aoede Gemini defaultBreezy, conversational.
CharonInformative, measured.
FenrirExcitable, energetic.
KoreFirm, assertive.
LedaYouthful, bright.
OrusFirm, deeper register.
PuckUpbeat, playful.
ZephyrBright, clear.

Leaving the dropdown on Provider default falls back to your org’s default provider’s default voice (Eve for xAI, alloy for OpenAI, Aoede for Gemini).

Troubleshooting