@pyai/twilio
Turn a phone number into a production voice agent — in one line of code.
@pyai/twilio bridges a live Twilio call to a PyAI Omni
voice agent. Point a Twilio number at a tiny server, hand the Media Streams
WebSocket to OmniAgent.bridge(...), and your caller is instantly talking to a
real listen → think → speak agent — with sub-500 ms turn-taking, natural
barge-in, DTMF, and live transfer-to-human. You write zero audio or DSP code.
import { OmniAgent } from "@pyai/twilio";
OmniAgent.bridge(twilioWebSocket, {
apiKey: process.env.PYAI_API_KEY!,
agentId: "support-bot",
voice: "stock_sarah_style2",
persona: "You are a warm, concise support agent for Acme.",
knowledge: async (q) => myVectorSearch(q),
});
That single call is a complete phone agent.
Why Omni — one SDK for the whole voice agent
Most voice stacks are a fragile chain of four vendors: speech-to-text → an LLM →
text-to-speech → a telephony bridge, each with its own latency, billing, and
failure mode. Omni collapses all of it into one realtime engine behind one
WebSocket and one API key:
- End-to-end, not glued together. Transcription, the reasoning brain,
retrieval, and synthesis run as a single speech-to-speech model — so you ship
an agent, not an integration project.
- Built for the phone. ~431 ms median turn-taking with real barge-in, tuned
for 8 kHz call audio. Conversations feel human, not walkie-talkie.
- Grounded and capable. Bind your own knowledge per turn and (roadmap) call
your tools, so the agent answers from your content and takes real actions.
- One predictable bill. Omni is all-in at $0.05/min, billed per second —
no four-vendor passthrough, no surprise invoice.
- Open and portable. Standard WebSocket, opaque API key, OpenAI-compatible
surfaces. Nothing proprietary to lock you in.
@pyai/twilio is the last mile: it makes Omni answer a real phone number.
🎁 Get $50 in free credit when you sign up
Create a key at console.pyai.com — email only,
no credit card, no sales call. That's roughly 1,000 minutes of live Omni
calls to build and test with, free. Your key works on every surface instantly.
What the bridge handles for you
The bridge owns the entire audio path so you never touch a codec or a resampler:
- Codec & rate. Twilio speaks G.711 mu‑law @ 8 kHz; Omni speaks
PCM16. The bridge transcodes both ways and resamples with a proper
anti‑aliased polyphase filter (not naive decimation) — 8 kHz ⇄ 24 kHz is a
clean 3:1, 8 kHz ⇄ 16 kHz is 2:1. Run Omni at
omniRate: 8000 to skip
resampling entirely.
- Handshake. Opens the Omni WebSocket with subprotocol auth and sends the
post‑handshake
configure frame (voice, persona, optional knowledge endpoint).
- Barge‑in. When the caller talks over the agent, Omni's barge‑in event is
relayed to Twilio as a
clear, so the agent stops mid‑word.
- DTMF. Caller key presses are forwarded to the agent.
- Knowledge. Your optional
knowledge(query) callback is invoked on each
finalized caller turn; whatever you return is pushed to the agent as grounding.
Install
npm install @pyai/twilio
Requires Node ≥ 22 (uses the ws WebSocket client; everything else is built‑in).
You'll need a PyAI key (pyai_live_… or a free pyai_test_… sandbox key) —
grab one with $50 free credit — and a Twilio
number with Media Streams.
1. TwiML: open a bidirectional stream
When Twilio receives a call it fetches TwiML from your webhook. Use
<Connect><Stream> (not <Start><Stream>) so the socket is two‑way and the
agent can speak back:
<?xml version="1.0" encoding="UTF-8"?>
<Response>
<Connect>
<Stream url="wss://your-host.example.com/media" />
</Connect>
</Response>
connectStreamTwiML("wss://your-host/media") builds exactly this string for you.
Set the number's A call comes in webhook to https://your-host/voice.
2. A ~10‑line Node server
import Fastify from "fastify";
import websocket from "@fastify/websocket";
import { OmniAgent, connectStreamTwiML } from "@pyai/twilio";
const app = Fastify();
await app.register(websocket);
app.post("/voice", (req, reply) =>
reply.type("text/xml").send(connectStreamTwiML(`wss://${req.headers.host}/media`)));
app.get("/media", { websocket: true }, (twilioWS) =>
OmniAgent.bridge(twilioWS, { apiKey: process.env.PYAI_API_KEY, agentId: "support-bot" }));
await app.listen({ port: 8080, host: "0.0.0.0" });
Tunnel it (ngrok http 8080), point your Twilio number at https://<host>/voice,
and call the number. See examples/twilio-omni-voice-agent
for a complete, runnable version.
API
OmniAgent.bridge(twilioWS, options) → BridgeHandle
twilioWS is the Twilio Media Streams socket (the Node ws socket your
framework hands you). Options:
apiKey | string | Required. pyai_live_… / pyai_test_…. Opaque — never parsed. |
agentId | string | Required. Opaque label authorized by your key's org. |
voice | string | Voice id (stock / clone / designed). |
persona | string | System prompt / role for the agent. |
knowledge | (q) => facts | Promise<facts> | Per‑turn grounding callback. |
kbEndpoint / kbToken | string | Customer‑hosted endpoint the engine pulls per turn. |
omniRate | 8000 | 16000 | 24000 | Omni session rate. Default 24000. 8000 = no resampling. |
baseURL | string | Defaults to https://api.pyai.com. |
onTranscript / onTransfer / onError / onClose | callbacks | Observability + lifecycle. |
Returns a handle with close() and the underlying omni client.
Lower‑level building blocks
Also exported for custom pipelines (and fully unit‑tested):
import {
muLawEncode, muLawDecode,
Resampler, makeResampler,
pcm16ToBytes, bytesToPcm16,
OmniClient, omniWsUrl,
parseTwilioMessage, twilioMedia,
twilioClear, connectStreamTwiML,
} from "@pyai/twilio";
Notes & limits
- Barge‑in / event names. A couple of Omni event field names aren't yet
byte‑pinned across the protocol doc and the Twilio guide; the client accepts
both spellings and isolates them in
src/omni.ts so they're trivial to update.
- Transfer to a human. The bridge surfaces Omni's transfer event via
onTransfer(info); perform the actual call redirect with Twilio's REST API
(see the example) — that needs your Twilio credentials, not PyAI's.
- Outbound framing. Agent audio is reframed into ~20 ms (160‑byte) mu‑law
frames tagged with the
streamSid, which is what Twilio's jitter buffer likes.
Develop
npm install
npm test
npm run typecheck
npm run build
MIT licensed.