Skip to content
Talk to us
For engineers · Architecture

The science under the hood. Latency, inference, compute, state.

How a spoken sentence becomes a typed decision, a model call, an action and a durable memory, and where every millisecond goes. Every figure on this page is cited to the published documentation of the infrastructure we build on.

~2ms1Agent to GPU. STT, LLM and TTS colocated with the agent runtime.
50–160ms9End-of-turn model, per turn. Context-aware, not just silence.
<100ms11To the nearest media server, anywhere. Mesh design target.
<200ms6Round trip with GPUs colocated at the Sydney PoP.
400+20Models routable through one API. 464 counted on Oct 6, 2026.
1Mtokens4Context on hosted open-weight models.
hot_path · inbound voiceschematic
  1. PSTNcaller
  2. SIPon-net
  3. SFU roomWebRTC
  4. Agent workerisolated~2 ms
  5. GPUSTT · LLM · TTS
Caller → PSTN → Telnyx SIP → LiveKit SIP bridge → Room → Agent worker 1
01Latency

Anatomy of a voice turn. Where the milliseconds go.

A reply feels human when the stages overlap. Streaming STT, a context-aware end-of-turn model, preemptive generation and colocated GPUs each remove a serial wait. Flip them to see the shape change.

trace · single turnschematic · not to scale except labelled values
caller.speech
stt.streaminterim → final · Nova-3 / Flux on GPU 5
turn.end_of_utterancemin delay 0.3 s · model ~50–160 ms/turn 9
net.agent→gpu~2 ms colocated 1
llm.generate (preemptive)starts on final transcript, before the turn is confirmed 10
tts.streamstreams as tokens arrive
agent.audio_out
Documented defaults and limits
ParameterValueEffect
endpointing.min_delay0.5 s · 0.3 s9Silence before a turn can end; lower with the audio turn detector.
endpointing.max_delay3.0 s10Upper bound when the model thinks the caller will continue.
turn_detector.latency~50–160 ms9Per-turn inference for the open-weights end-of-turn model, 396 MB, 14 languages.
interruption.min_duration0.5 s10Speech needed before the agent yields; adaptive mode ignores backchannels like “mm-hm”.
resume_false_interruption2.0 s10If no transcript follows a barge-in, the agent resumes where it stopped.
agent → gpu~2 ms1STT, LLM and TTS colocated with the agent runtime. No external API on the hot path.
02Inference

Classify first. Then spend compute where it pays.

Most decisions in a business are bounded: route, score, qualify, escalate. Those go to a non-generative typed model. Generation is reserved for the work that needs it, on the lane built for it.

request
lane 1

Typed decision

Jev · non-generative

Choice, Score and Noul questions evaluated in parallel against one state. Returns typed values, probabilities and confidence 18. Published example: 0.114 s vs 8.566 s for an LLM 19.

lane 2

Voice LLM

real-time · reasoning off

Voice-verified open-weight models on colocated GPUs. Reasoning is always disabled on voice calls 4.

lane 3

Long context

up to 1M tokens

Hosted open-weight models with 1M-token context for whole-history summaries and research 4.

lane 4

Model router

400+ models

Any model in a 400+ catalogue behind one OpenAI-compatible API, picked per task for quality, cost and latency 20.

→ lane 1 · Typed decision

A bounded question with a fixed answer set. No prose needed.

score: 0..1 · choice: "call_now" | "nurture" | "disqualify" · probabilities · confidenceIllustrative routing. Thresholds and model choices are configured per workspace.
Hosted open-weight models on colocated GPUs (listed by Telnyx, Oct 2026)4
ModelParamsContextVoice
moonshotai/Kimi-K2.61.0T256KVoice default
zai-org/GLM-5.2753.9B1MVoice verified
moonshotai/Kimi-K32.8T1M—
zai-org/GLM-5.3753.9B1M—
MiniMaxAI/MiniMax-M3-MXFP8428B1M—
Qwen/Qwen3.8-27B27B256K—

Catalogues change often. Embeddings: Qwen3-Embedding-8B, 4,096 dimensions, 8,192-token input 4.

03Compute

GPUs beside the phone network. Every region a full stack.

Inference runs in the same facilities that terminate the call. Each voice region carries its own SFU, SIP, agent runtime and models, so the hot path never leaves the building.

schematic positions
5 · 4GPU regions · continents for inference3
17+Telephony PoPs on a private IP backbone7
failoverTo the next-lowest-latency region, not an error3
statelessChat-completions requests and responses are not stored3
04State and memory

Stateful by construction. One actor per customer.

Every event for a customer routes to the same durable actor. Requests are serialized, so two agents never race on one record, and state recovers after a restart. Context for each turn is assembled from facts, recent history and retrieval.

01 ingestcall · sms · email · web · review
02 routeper-entity, one mailbox
03 persistdurable writes, merge-patch state
04 indexembeddings, 4,096-d
05 assemblefacts + recent + retrieved
customer_actor.tsillustrative shape
type CustomerActor = {
  id: CustomerId                 // routes every event for this customer here
  timeline: Event[]              // calls, texts, email, visits, reviews
  facts: Fact[]                  // each with source + observed_at
  tasks: ScheduledTask[]         // durable, survive restarts
  policy: ContactPolicy          // who may contact, when, on which channel
}

// One mailbox per customer: requests are serialized, writes are durable.
on(event)  → append(timeline) → extract(facts) → embed(chunks)
on(turn)   → assemble(context: facts + recent + retrieved) → model

Building blocks: per-entity routing, serialized requests, durable writes and recovery 16; persistent history, durable scheduled tasks and merge-patch state 17; embeddings 4.

05Real-time media

WebRTC at the edge. People can step in mid-call.

Distributed SFU mesh

Each participant connects to the nearest SFU; servers relay tracks to each other over a FlatBuffers-based protocol carrying RTP. Design target: a media server within 100 ms of anyone, 99.99+% availability 11.

Supervisor as a third leg

A manager joins a live call as a third leg without converting it to a conference, so the agent and caller are undisturbed 14.

Warm transfer with context

Hold the caller, brief the person privately, then connect. If nobody picks up, the call returns to the original flow 13.

Turn-aware STT

Deepgram Flux on colocated GPUs emits end-of-turn signals from the transcription stream itself, with tunable eager thresholds 5.

06Observability and evals

Every turn is a trace. Every version is scored.

Audio, transcripts, traces and logs share one session timeline. Assistant versions are compared on instruction-following and estimated caller satisfaction before they ship wider.

session timelinetrace shape · illustrative · no measured values
  1. session
  2. turn[3]
  3. vad.speech
  4. stt.final
  5. turn.eou
  6. jev.decide intent
  7. tool.crm.lookup
  8. llm.generate
  9. tts.stream
  10. eval.instruction_following
  11. memory.write facts

Session timeline 12; managed insights and version comparison 15.

07One endpoint

Natural language in. Typed plans and approvals out.

One orchestration endpoint turns a sentence into a plan, typed decisions and actions. Anything above the workspace’s autonomy level stops for approval, and every step lands in memory.

illustrative shape · not a published API
POST /v1/ask
Content-Type: application/json

{
  "workspace": "ws_demo",
  "input": "Call back everyone who asked about pricing this week.",
  "autonomy": "with_approval",
  "stream": true
}
08Sources

Check our work. Every number, cited.

Figures are vendor-published numbers, defaults and design targets for the infrastructure IntelAgents builds on, checked on October 6, 2026. They are not IntelAgents benchmarks; real latency depends on region, models, carrier and network.

  1. [1]
    Telnyx · LiveKit on Telnyx: Architecture

    STT, TTS and LLM on GPUs colocated with the agent runtime, approximately 2 ms from the agent; on-net inbound call flow; autoscaled, isolated agent workers; each region a full stack.

  2. [2]
    Telnyx · LiveKit on Telnyx: Regions

    Platform, agent and SIP endpoints in New York, San Francisco, Atlanta and Sydney.

  3. [3]
    Telnyx · Inference: Regions and availability

    GPU infrastructure across five regions on four continents; latency-based routing by ingress domain; failover to the next-lowest-latency region; storage governed by data locality.

  4. [4]
    Telnyx · Inference: Available models

    Hosted open-weight models with parameter counts, context lengths and voice verification; reasoning disabled on voice calls; embedding models and dimensions.

  5. [5]
    Telnyx · LiveKit on Telnyx: STT models

    Deepgram Nova-3, Nova-2 and Flux hosted on Telnyx GPUs; Flux has built-in end-of-turn detection.

  6. [6]
    Telnyx · Sydney GPU deployment (release note, Oct 2025)

    GPUs colocated with the Sydney telephony PoP; round-trip time under 200 ms.

  7. [7]
    Telnyx · Telnyx deploys GPUs in Sydney (press release, Oct 29, 2025)

    17+ telephony PoPs on a private IP backbone; GPUs in North America, Europe and APAC.

  8. [8]
    Telnyx · LiveKit on Telnyx: Overview

    Built-in SIP with HD voice (G.722 and Opus).

  9. [9]
    LiveKit · Turn detector

    Open-weights end-of-turn model: ~50–160 ms per turn, 396 MB on disk, 14 languages; endpointing minimum delay 0.5 s by default, 0.3 s with the audio detector.

  10. [10]
    LiveKit · Turn-taking tuning

    Preemptive generation; adaptive interruption handling; false-interruption recovery after 2.0 s; interruption minimum 0.5 s; endpointing maximum 3.0 s.

  11. [11]
    LiveKit · Scaling WebRTC with distributed mesh

    A media server within 100 ms of anyone; SFU-to-SFU relay over a FlatBuffers-based protocol with RTP; 99.99+% availability target.

  12. [12]
    LiveKit · Agent Insights

    Audio, transcripts, traces and logs on one session timeline.

  13. [13]
    LiveKit · Warm transfer

    Hold the caller, consult privately, pass context, then connect the person.

  14. [14]
    Telnyx · Supervising leg

    A supervisor joins a live call as a third leg without converting it to a conference.

  15. [15]
    Telnyx · Managed insights

    Instruction-following and estimated caller-satisfaction rubrics, compared across assistant versions.

  16. [16]
    Telnyx · Stateful Actors

    Per-entity routing, serialized requests, durable writes and recovery after restart.

  17. [17]
    Telnyx · AgentSDK

    Persistent message history, durable scheduled tasks and merge-patch state.

  18. [18]
    TypeSafe · Introduction to Jev

    Choice, Score and Noul questions evaluated in parallel and in isolation against one state; typed answers with probabilities and confidence.

  19. [19]
    TypeSafe · typesafe.ai

    Published workflow example completed in 0.114 s, against 8.566 s for an LLM on the same task.

  20. [20]
    OpenRouter · Models API

    464 models listed when counted on October 6, 2026.

09Next

Bring your architects. We’ll bring the diagrams.

Walk through your regions, carriers, models and data residency with our team.