Typed decision
Jev · non-generativeChoice, Score and Noul questions evaluated in parallel against one state. Returns typed values, probabilities and confidence 18. Published example: 0.114 s vs 8.566 s for an LLM 19.
Talk to usHow a spoken sentence becomes a typed decision, a model call, an action and a durable memory, and where every millisecond goes. Every figure on this page is cited to the published documentation of the infrastructure we build on.
A reply feels human when the stages overlap. Streaming STT, a context-aware end-of-turn model, preemptive generation and colocated GPUs each remove a serial wait. Flip them to see the shape change.
| Parameter | Value | Effect |
|---|---|---|
| endpointing.min_delay | 0.5 s · 0.3 s9 | Silence before a turn can end; lower with the audio turn detector. |
| endpointing.max_delay | 3.0 s10 | Upper bound when the model thinks the caller will continue. |
| turn_detector.latency | ~50–160 ms9 | Per-turn inference for the open-weights end-of-turn model, 396 MB, 14 languages. |
| interruption.min_duration | 0.5 s10 | Speech needed before the agent yields; adaptive mode ignores backchannels like “mm-hm”. |
| resume_false_interruption | 2.0 s10 | If no transcript follows a barge-in, the agent resumes where it stopped. |
| agent → gpu | ~2 ms1 | STT, LLM and TTS colocated with the agent runtime. No external API on the hot path. |
Most decisions in a business are bounded: route, score, qualify, escalate. Those go to a non-generative typed model. Generation is reserved for the work that needs it, on the lane built for it.
Choice, Score and Noul questions evaluated in parallel against one state. Returns typed values, probabilities and confidence 18. Published example: 0.114 s vs 8.566 s for an LLM 19.
Voice-verified open-weight models on colocated GPUs. Reasoning is always disabled on voice calls 4.
Hosted open-weight models with 1M-token context for whole-history summaries and research 4.
Any model in a 400+ catalogue behind one OpenAI-compatible API, picked per task for quality, cost and latency 20.
A bounded question with a fixed answer set. No prose needed.
score: 0..1 · choice: "call_now" | "nurture" | "disqualify" · probabilities · confidenceIllustrative routing. Thresholds and model choices are configured per workspace.| Model | Params | Context | Voice |
|---|---|---|---|
| moonshotai/Kimi-K2.6 | 1.0T | 256K | Voice default |
| zai-org/GLM-5.2 | 753.9B | 1M | Voice verified |
| moonshotai/Kimi-K3 | 2.8T | 1M | — |
| zai-org/GLM-5.3 | 753.9B | 1M | — |
| MiniMaxAI/MiniMax-M3-MXFP8 | 428B | 1M | — |
| Qwen/Qwen3.8-27B | 27B | 256K | — |
Catalogues change often. Embeddings: Qwen3-Embedding-8B, 4,096 dimensions, 8,192-token input 4.
Inference runs in the same facilities that terminate the call. Each voice region carries its own SFU, SIP, agent runtime and models, so the hot path never leaves the building.
Every event for a customer routes to the same durable actor. Requests are serialized, so two agents never race on one record, and state recovers after a restart. Context for each turn is assembled from facts, recent history and retrieval.
type CustomerActor = {
id: CustomerId // routes every event for this customer here
timeline: Event[] // calls, texts, email, visits, reviews
facts: Fact[] // each with source + observed_at
tasks: ScheduledTask[] // durable, survive restarts
policy: ContactPolicy // who may contact, when, on which channel
}
// One mailbox per customer: requests are serialized, writes are durable.
on(event) → append(timeline) → extract(facts) → embed(chunks)
on(turn) → assemble(context: facts + recent + retrieved) → modelBuilding blocks: per-entity routing, serialized requests, durable writes and recovery 16; persistent history, durable scheduled tasks and merge-patch state 17; embeddings 4.
Each participant connects to the nearest SFU; servers relay tracks to each other over a FlatBuffers-based protocol carrying RTP. Design target: a media server within 100 ms of anyone, 99.99+% availability 11.
A manager joins a live call as a third leg without converting it to a conference, so the agent and caller are undisturbed 14.
Hold the caller, brief the person privately, then connect. If nobody picks up, the call returns to the original flow 13.
Deepgram Flux on colocated GPUs emits end-of-turn signals from the transcription stream itself, with tunable eager thresholds 5.
Audio, transcripts, traces and logs share one session timeline. Assistant versions are compared on instruction-following and estimated caller satisfaction before they ship wider.
Session timeline 12; managed insights and version comparison 15.
One orchestration endpoint turns a sentence into a plan, typed decisions and actions. Anything above the workspace’s autonomy level stops for approval, and every step lands in memory.
POST /v1/ask
Content-Type: application/json
{
"workspace": "ws_demo",
"input": "Call back everyone who asked about pricing this week.",
"autonomy": "with_approval",
"stream": true
}Figures are vendor-published numbers, defaults and design targets for the infrastructure IntelAgents builds on, checked on October 6, 2026. They are not IntelAgents benchmarks; real latency depends on region, models, carrier and network.
STT, TTS and LLM on GPUs colocated with the agent runtime, approximately 2 ms from the agent; on-net inbound call flow; autoscaled, isolated agent workers; each region a full stack.
Platform, agent and SIP endpoints in New York, San Francisco, Atlanta and Sydney.
GPU infrastructure across five regions on four continents; latency-based routing by ingress domain; failover to the next-lowest-latency region; storage governed by data locality.
Hosted open-weight models with parameter counts, context lengths and voice verification; reasoning disabled on voice calls; embedding models and dimensions.
Deepgram Nova-3, Nova-2 and Flux hosted on Telnyx GPUs; Flux has built-in end-of-turn detection.
GPUs colocated with the Sydney telephony PoP; round-trip time under 200 ms.
17+ telephony PoPs on a private IP backbone; GPUs in North America, Europe and APAC.
Built-in SIP with HD voice (G.722 and Opus).
Open-weights end-of-turn model: ~50–160 ms per turn, 396 MB on disk, 14 languages; endpointing minimum delay 0.5 s by default, 0.3 s with the audio detector.
Preemptive generation; adaptive interruption handling; false-interruption recovery after 2.0 s; interruption minimum 0.5 s; endpointing maximum 3.0 s.
A media server within 100 ms of anyone; SFU-to-SFU relay over a FlatBuffers-based protocol with RTP; 99.99+% availability target.
Audio, transcripts, traces and logs on one session timeline.
Hold the caller, consult privately, pass context, then connect the person.
A supervisor joins a live call as a third leg without converting it to a conference.
Instruction-following and estimated caller-satisfaction rubrics, compared across assistant versions.
Per-entity routing, serialized requests, durable writes and recovery after restart.
Persistent message history, durable scheduled tasks and merge-patch state.
Choice, Score and Noul questions evaluated in parallel and in isolation against one state; typed answers with probabilities and confidence.
Published workflow example completed in 0.114 s, against 8.566 s for an LLM on the same task.
464 models listed when counted on October 6, 2026.
Walk through your regions, carriers, models and data residency with our team.