Configure
Voice chat
Visitors can switch to a spoken conversation with your agent at any point. A single tap on the headphone chip starts a real-time voice call through xAI’s Grok voices. Interruption, transcript capture, and graceful fallbacks to text are all handled for you.
Turn it on
Almost all of setup lives on one screen: Settings → Voice. In the figure below, the Call behavior card sets the Silence timeout and Max call length and picks a Default voice provider; the Voice API keys card underneath is where each provider either Reuses your text-chat key or takes a separate voice key — a green Ready badge appears the moment a usable key is on file. Only the per-widget Enable voice chat toggle and layout live elsewhere (on the Customize page — see below).
Call behavior
Defaults applied to every voice session — changes take effect on the next call.
Voice API keys
Voice keys are encrypted at rest. A Ready badge means a usable key is on file — your own, or the platform-default xAI key.
The Settings → Voice tab — call behavior on top, per-provider keys below; Ready confirms a usable key.
- Open Settings → Voice in the dashboard.
- Pick your provider (xAI Grok, OpenAI Realtime, or Google Gemini Live) and how to authenticate. The simplest option is Reuse my text-chat key — if text chat already works with that provider, voice just works too. Otherwise paste a voice-enabled key (console.x.ai for xAI, platform.openai.com for OpenAI, aistudio.google.com for Gemini).
- Pick a silence timeout (default 10s) and max call length (default 15 min).
- Open Billing → Voice and pick a plan (see below). Metered customers get minute bundles; BYOK flat gets unlimited minutes for a flat $49/mo platform fee.
- In Customize & Deploy → Appearance, pick a voice UI layout and flip the Enable voice chat switch for each widget you want the chip to appear on. Voice stays hidden on widgets where the toggle is off.
- Click Try voice live on the same page to test the selected layout and voice end-to-end in an embedded preview before visitors ever see it.
You don’t need your own xAI key to try voice — Quincer AI ships a platform-default key that works out of the box on every plan. Pasting your own only matters if you want to bill xAI directly (BYOK) or use a region-specific key.
Pricing
Two billing models, both requiring you to bring your own xAI API key:
| Model | Minutes | Price | Best for |
|---|---|---|---|
| Voice 60 Monthly | 60/mo | $26/mo | Trying voice on a single widget |
| Voice 180 Monthly | 180/mo | $59/mo | Moderate use across a couple of widgets |
| Voice 960 Monthly | 960/mo | $249/mo | High-volume support / lead teams |
| Voice 60 / 180 / 960 Add-on | one-time | $29 / $69 / $299 | Top up a busy period without changing your plan |
| Voice BYOK (unlimited) | Unlimited | $49/mo | You pay xAI directly for usage; we charge a flat platform fee |
Monthly plans are ~10–17% cheaper than the equivalent one-time add-on.
UI layouts
Voice is visually distinct from text, so Quincer AI offers four layouts for how the call is presented inside the widget. Both live on the Customize & Deploy page: the Enable voice chat toggle is on the AI agent tab (it needs a voice plan on the widget’s owner), and once it’s on, the Voice UI layout card appears with the four-way picker below. In the figure, the dropdown carries the exact labels the product shows — Hybrid (recommended), Takeover, Inline composer, and Compact strip — and a live preview under it renders the chosen layout before you save. The Try voice live button (top-right of the preview panel) opens the real widget so you can test the layout and voice end-to-end.
Voice UI layout
How voice appears when a visitor activates it. Mobile widths auto-promote Hybrid to Takeover.
Also: Takeover · Inline composer · Compact strip. A live preview renders the chosen layout below the dropdown.
On Customize & Deploy — enable voice on the AI agent tab, then pick one of four layouts; Try voice live tests it end-to-end.
A live preview shows exactly how your chosen layout looks before you save.
| Layout | What it looks like | Best for |
|---|---|---|
| Hybrid default | A compact voice strip slides in above the chat transcript. Transcripts for both sides render inline in the message list as italic bubbles with a mic glyph, so visitors can scroll back through what was said. | Most widgets. Works on desktop and mobile; auto-promotes to takeover under 480px. |
| Takeover | The whole chat panel swaps to an orb-centric voice UI while the call is live. Text history is hidden until the call ends. Maximum focus. | Support flows where voice is the primary channel, or mobile-first deployments. |
| Inline | No dedicated strip or orb — transcripts simply appear in the normal message list while a mic-state dot in the header indicates listening / thinking / speaking. | Minimal UIs that want voice to feel like dictation rather than a “call”. |
| Strip | A slim always-visible control strip at the top of the chat when a call is active. Transcripts render in the message list like Inline, but with richer on-strip controls. | Teams that want the call state obvious without taking over the whole panel. |
On screens narrower than 480px, Hybrid and Strip automatically promote to Takeover — there’s not enough room to show a transcript and voice controls side-by-side. The dashboard preview has a Desktop/Mobile toggle so you can see both states.
How it works
- Chip in the widget header. When voice is enabled, a small headphone icon appears next to the close button. Visitors can also type “call me” or “switch to voice” to trigger the same flow.
-
Server-side relay. All voice audio is brokered through our dedicated
relay service at
voice.quincer.com. Your xAI API key never touches the browser; the relay authenticates with xAI and streams audio both ways. - Knowledge-grounded answers. On every user utterance the relay queries your knowledge base — the same cross-persona vector search text chat uses — and injects the top matches into the model’s context before it speaks. You maintain one KB; both channels stay on the same facts.
- Real-time barge-in. The moment the visitor starts speaking, the bot stops talking. No waiting, no hold-to-talk. The one exception is your opening prompt, which is protected by default — see Interrupting the opening prompt.
- Integration tools. Every tool enabled on a persona (HubSpot, Pipeline CRM, Calendly, Gmail, etc.) is callable from voice too. Ask the agent to book a meeting, create a contact, or log a ticket — it narrates what it just did.
- Silence + max-length protection. If the agent hears nothing for your configured silence window (default 10s), it politely winds down and switches back to text. Same behavior when it hits the max-call-length safeguard (default 15 min).
- Graceful cap handling. Metered customers who run out of minutes mid-call hear a natural wind-down line (“I need to switch back to text as my voice time has run out”) rather than a silent hang-up.
- 80% warning email. Metered customers get a one-time email per period when they cross 80% usage, with a one-click link to buy more minutes.
-
Transcripts persist to the inbox. Every voice turn — visitor and
assistant — is saved as a message on the same
Conversationrow a text chat would use, so the live inbox, lead timeline, and Slack escalations all see voice and text as one unified conversation.
Interrupting the opening prompt
When a call connects, your agent speaks an opening prompt. While that prompt is playing the caller’s microphone is ignored, so a door closing, a colleague talking nearby, or hold music bleeding in can’t cut the greeting off half way through. This is on by default and it is what we recommend.
It matters more than it sounds. Without it, the speech detector can trigger on the greeting itself or on background noise, and the call starts mid-sentence with the agent already listening to nothing. On one live call that misfire was transcribed as a different language and locked the whole conversation to it.
You’ll find the control in Customize → Deploy, directly under Enable voice chat:
- Mic off on first prompt — Recommended (ticked by default). The mic stays off until the opening prompt finishes.
- Untick it and you can set Open the mic after N seconds. The caller can then interrupt after that long, even if the prompt is still playing. Minimum 3 seconds.
You don’t need to know how long your prompt is. The setting can only make the mic open earlier — never later. If your prompt finishes before the time you set, it finishes first and your setting simply has no effect. Leave the seconds field blank to keep the mic off for the whole prompt.
Why a 3-second minimum? Below that, ordinary background noise interrupts the prompt before the caller has heard enough of it to know who they’ve reached — which is the problem the protection exists to solve.
Unticking the box without setting a number means barge-in from the very first millisecond. That has always been available and is unchanged; the seconds field is the middle ground between the two.
Live voice takeover
When an AI voice call needs a human, any licensed teammate can take it over mid-conversation. Visitor voice stays on Quincer AI’s relay while the AI is talking; the moment an agent clicks Take Over on the Live Conversations page, the widget bridges the visitor onto Cloudflare’s WebRTC SFU and the agent joins the same call. No app install for the visitor, no phone number, no third-party dial-in.
In the figure, the open conversation carries a Voice call badge in its header while the AI is live, and the action bar at the bottom shows the single Take Over button. If you don’t hold a live-voice seat, clicking it surfaces the amber gate notice shown here instead of the call console.
A live voice call on Live Conversations — the Voice call badge, one Take Over button, and the seat-gate notice.
- On Live Conversations the active voice call shows a Voice call badge. Open the conversation.
- Click Take Over. If you have a live-voice seat (see below), you’re redirected to the call console; if not, you get a banner suggesting the visitor be asked to switch to text.
- The AI speaks a short hand-off line (“Stand by — Mo is joining the call now”) while your console auto-requests mic access.
- Within a few seconds, your console and the visitor’s widget are bridged directly via Cloudflare Realtime. The AI steps aside. You and the visitor hear each other in real time.
- Either side clicking Hang up (or the visitor clicking End call) ends the bridge and closes the call for the other party.
If the WebRTC bridge fails (for example, firewalled networks), the AI stays on the call and tells the visitor: “I’m getting an error transferring you — let me stay on and help out, or switch to text and my colleague will chat with you there.” The session continues; no dropped call.
Seats & licensing
Live voice takeover is a per-seat SKU at $19/user/month. Assign seats on Organization → People, using the Voice seat toggle on each teammate’s row. Super-admins can also grant seats directly from Admin → Users for free while you’re setting up. Visitor voice minutes continue to bill from the org’s voice plan; the $19 seat covers only the human handover side.
Keypad entry (touch-tones)
On phone calls, callers can answer with the keypad instead of their voice — an account number, an order number, a PIN. Your persona receives the digits as caller input, answers them like any spoken turn, and the entry is saved to the conversation transcript so you can see exactly what was keyed.
An entry is complete when any of these happens:
-
The caller presses
#— the entry is submitted immediately. The#itself is a terminator, not part of the entry. - The caller stops typing for about 2.5 seconds — long enough that reading a number off a bill doesn’t get split in two, short enough that the conversation keeps moving.
- The entry reaches 32 digits — it is submitted rather than truncated, so the persona never sees a shortened number that looks complete.
A caller who is mid-entry is a caller who is present: typing resets the silence timeout, so keying a long number never trips the “are you still there?” wind-down. And if the persona is mid-sentence when an entry lands, it stops talking and answers the entry.
Press 0 for a person
A lone 0 followed by that same ~2.5-second pause is not data — it’s the universal “get me a human”, and it triggers the same transfer-to-a-human flow as asking out loud: a warm transfer to the persona’s configured transfer target, or a team notification when none is set. One escalation per call — repeated 0 presses after the first are ignored.
A 0 inside a longer entry is just a zero. Account numbers and phone numbers
contain zeros routinely; entering 5501# sends the persona
5501 — it never transfers the call.
The caller pressed 4 8 2 1 # — # submitted the entry immediately, and the transcript records exactly what the persona received.
On Google Gemini Live, keypad entries can’t be delivered to the model (its protocol has no channel for them) — callers can still speak the digits. Pressing a lone 0 to reach a person works on every provider: it’s handled before the AI ever sees it.
Voice providers
Quincer AI ships with three production voice providers — xAI Grok, OpenAI Realtime, and Google Gemini Live. Pick a default provider under Settings → Voice → Default voice provider; individual personas can select any voice from any enabled provider. All three run over the same relay pipeline, so everything else (layouts, billing, transcripts, knowledge grounding, integration tools) behaves identically.
| Provider | Model | Good for |
|---|---|---|
| xAI Grok | grok-voice-think-fast-1.0 default · grok-voice-fast-1.0 | Fast, conversational, five curated voices. Grok Think Fast is the flagship — background reasoning, stronger tool-calling, and 20+ languages including Arabic (Egypt, Saudi Arabia, UAE), Bengali, Hindi, Korean, Turkish, Vietnamese, and more. Legacy Grok Fast is still selectable in Settings → Voice while xAI completes its migration. |
| OpenAI Realtime | gpt-realtime | Ten voices including the expressive marin and cedar. Strongest when you need nuanced delivery. |
| Google Gemini Live | gemini-3.1-flash-live | Eight voices (Aoede, Charon, Fenrir, Kore, Leda, Orus, Puck, Zephyr). 70+ languages out of the box, strong multilingual delivery. |
Pick the OpenAI realtime model
Running voice on OpenAI? The model is yours to choose. When your
Default voice provider is OpenAI Realtime (gpt-realtime),
an OpenAI realtime model field appears directly beneath the provider picker
on Settings → Voice. Type any realtime model id — including a new
release the day it ships. Leave it blank to stay on the platform default
(gpt-realtime); the help text under the field always shows which model calls
currently use.
The OpenAI realtime model field on Settings → Voice — it appears once the default provider is OpenAI Realtime; blank means the platform default.
Two things to know before you type a custom id. It must be a realtime speech-to-speech model — a model id OpenAI doesn’t recognise fails when the call connects, not when you save. And billing follows the model: voice cost is looked up by model id, so the model you pick is the model that answers and the one your usage reflects. Any custom id shows a Check pricing for this model reminder: if we hold no price for it yet, calls still work but their cost reports as $0 until a price is added. For the full model catalogue and how model pricing works across the product, see AI models.
Voices
Each persona picks one voice on its Voice & phone card, under Personas → [Persona]. In the figure, the Voice (optional) dropdown groups voices by provider (xAI, OpenAI Realtime, Gemini Live) and the first option is Provider default; the Preview button beside it plays a short sample so you can compare tone before saving. If you pick a voice from a provider that isn’t your workspace’s active one, an inline Mismatch warning appears with a link to Settings → Voice — Preview still works so you can audition any voice.
The call voice, spoken phone greeting, and default call language for this persona.
Groups: xAI (Grok Voice) · OpenAI Realtime · Google Gemini Live. First option is Provider default (Eve / alloy / Aoede).
A persona’s Voice & phone card — a provider-grouped dropdown plus a Preview button to audition tone before saving.
xAI Grok
| Voice | Tone |
|---|---|
| Eve xAI default | Warm, balanced female voice. Good all-rounder. |
| Ara | Softer female, calmer cadence. |
| Leo | Friendly male, conversational. |
| Rex | Deeper male, more confident delivery. |
| Sal | Neutral, professional — works well for support. |
OpenAI Realtime
| Voice | Tone |
|---|---|
| alloy OpenAI default | Balanced, versatile. OpenAI’s neutral-tone pick. |
| marin recommended | Most natural delivery. Recommended for nuanced conversations. |
| cedar recommended | Most natural delivery. Slightly richer than marin. |
| ash | Expressive, natural male. |
| ballad | Warm, storytelling. |
| coral | Friendly, upbeat female. |
| echo | Clear, confident male. |
| sage | Calm, thoughtful. |
| shimmer | Bright, energetic female. |
| verse | Smooth, musical. |
Google Gemini Live
| Voice | Tone |
|---|---|
| Aoede Gemini default | Breezy, conversational. |
| Charon | Informative, measured. |
| Fenrir | Excitable, energetic. |
| Kore | Firm, assertive. |
| Leda | Youthful, bright. |
| Orus | Firm, deeper register. |
| Puck | Upbeat, playful. |
| Zephyr | Bright, clear. |
Leaving the dropdown on Provider default falls back to your org’s default provider’s default voice (Eve for xAI, alloy for OpenAI, Aoede for Gemini).
Troubleshooting
- Chip not showing. Confirm Enable voice chat is on for the current widget in Customize & Deploy, and that you’re on a voice plan. The chip is hidden per-widget so agencies can keep voice off on some deployments.
- “Voice is not set up for this widget”. No voice API key resolvable. The Voice tab in Settings shows a Ready badge when either your own key is on file or the platform-default xAI key is available. If nothing is ready, paste an xAI key under Enter a separate voice key.
- “You’ve used all your voice minutes”. Metered bundle exhausted this period. Buy an add-on on the Billing → Voice page or wait for the next reset. If you think your plan should include more minutes, contact support — comp minutes can be granted manually per org.
- Agent invents things about your company. Voice is grounded in your knowledge base the same way text is, so missing facts usually mean the KB has no entry covering that topic yet. Add an “About” entry in Knowledge base describing what your company does, and re-test. Voice sessions always prime with an about/products summary at connect time, so improvements take effect on the next call.
- Long pause after the greeting. Fixed platform-wide — words a caller spoke during or right after the greeting used to be dropped, leaving dead air before the first answer. They now buffer into the conversation and the agent takes the turn as soon as the greeting ends. The fix is live on every voice-enabled brand; nothing to change on your side.
- Microphone blocked. The browser denied mic access. The visitor needs to grant permission in their browser settings and try again.
- Mobile layout feels cramped. Hybrid and Strip auto-promote to Takeover under 480px. If the layout looks wrong at desktop widths, use the Desktop/Mobile toggle inside Try voice live on the Customize page to inspect both.