— / live demo / a voice agent / five minutes a call
I can’t take every call. This half of me can.
This is a voice agent that answers as me. Press start, allow the microphone, and ask about my work the way you would on a phone screen. It only knows what this site says about me, and when it doesn’t know something it says so instead of making it up. It’s an AI, and it will tell you that itself.
a stock voice for now, not mine · the call isn’t recorded · you can type if you’d rather not talk
half adi / bars out: it’s talking / bars in: you are5:00 a call
Ready when you are.
Press start and say hello. Your browser will ask for the microphone; say no and you can still type to it.
- last reply
- —
- replies
- 0
- microphone
- —
Talk over it whenever you like. It stops.
what was saidthe models in use show here once it picks up
Nothing yet. The whole conversation is written here as it happens.
how long each reply took / measured by the agent on this call, not typed in
Nothing yet. After each reply, the agent reports how long it spent on each step and the row appears here.
02 / what happens when you talk
Six steps, and the hard one is knowing when you’ve stopped.
01 / Clean up the sound
Background noise is taken out first: a fan, a street, somebody else's meeting. Everything after this works on your voice alone.
02 / Notice you're speaking
A small model listens for speech, so nothing is spent on silence and it knows the moment you start talking over it.
03 / Wait until you've finished
People pause in the middle of a sentence. A second model reads what you've said so far and judges whether it's a finished thought, so it doesn't jump in at every breath.
04 / Write it down
Your speech becomes text as you say it. That text is the caption you see on the page.
05 / Work out a reply
A language model answers with two things in front of it: how I talk, and the facts from this site. It has all of them to hand, so it never stops to look something up.
06 / Say it
The reply is turned into speech and starts playing while the rest is still being written. Talk over it and it stops.
03 / what stops it making things up
It can only say what this site already says.
A voice that sounds like a person and invents their history is worse than no voice at all. So everything it knows about me is exported from this site’s own content by a script, and it’s told to answer from that and nothing else. It won’t quote a salary, promise a start date or guess at anything that’s mine to answer. It puts my contact details on your screen instead.
How it’s built, for anyone who wants it
The agent is a Python program on LiveKit’s agent framework. Your browser and the agent meet in a private room and exchange audio over WebRTC, which is built for live sound: a lost packet is skipped rather than waited for, so one bad moment on the network costs a blip instead of a stall.
The language model and the voice each have a backup behind them. If the first fails mid-call, the next one takes over and the call carries on. The models in use are shown above the transcript once a call connects, reported by the agent rather than written into this page.
The language model starts drafting the moment your words are transcribed, before it’s certain you’ve finished, and only speaks once that’s confirmed. It’s a reasonable trade here because nothing the agent does can’t be undone; it would be the wrong one for an agent that moves money.
After every reply the agent sends this page its own timings for that reply, which is the table under the call. Each call is capped at five minutes and each address at a few calls every ten minutes, because every second of it costs something.
04 / how fast it answers
From 2.25 seconds of silence to about 1.3, by measuring.
A pause that’s fine in a chat window feels broken on a call. So I built a test caller: a program that rings the agent, plays recorded questions into the line like a microphone would, and times the silence before it hears an answer. The first version left a gap of 2.25 seconds. The one you’re talking to typically leaves 1.3, and it isn’t under a second yet. I tried.
1.30 s
typical silence before a reply
median of 15 spoken questions
1.05 s
the quickest of the fifteen
6 of 15 came in under 1.2 s
2.53 s
the slowest of the fifteen
the language model stalled
0 of 1
times it cut in on a pause
a 1.1 s gap mid-sentence
measured 10 October 2026, with real speech, against the deployed agent
Where the second went
Looking things up cost half a second each time.
The first design kept a short summary in the model’s head and fetched detail on request. Every fetch was a second trip to the model before the first word. Now it holds everything, and a project’s link is put on your screen by plain code after the reply, not by the model.
One ordinary question was left hanging for three seconds.
The model that judges whether you’ve finished was only 47% sure about “So, what do you actually do?”, just under its cutoff, so it waited the maximum. A real mid-sentence pause never scored above 3%. Moving the cutoff into that gap fixed the question without making it interrupt.
The voice was hiding a quarter of a second.
One voice began every reply with 0.25 seconds of silence inside its own audio. No dashboard reports that, since the audio had technically started. It showed up only by measuring the sound itself. A different voice is audible 0.3 seconds sooner.
Saying “hmm” first made it worse.
The obvious trick is a quick “okay” the moment you stop, while the answer is prepared. As a separate word it got a sound out in 0.8 seconds and pushed the actual answer back by 0.6. As the first word of the reply it garbled the answers. It isn’t in use. The answer starting sooner is the only version of fast I’m willing to count.
Three things that sounded better and measured worse.
Two speech-to-speech models, the kind that skip transcription altogether: 2.04 and 1.58 seconds here, and one talked over the pause. An expressive mode that adds breaths and laughs: 0.3 to 0.4 seconds a reply. None is in use.
The model was chosen for not inventing things.
Several language models were equally quick. One newer one gave me a hobby I don’t have and got my job dates wrong. The one in use is from the family that was as fast as any and stuck to the facts.
What’s left is mostly one step.
Of the 1.3 seconds, about half a second is the language model getting to its first word, and on a slow moment that alone has been 1.8. Everything else is already near its floor.
05 / what it doesn’t do yet
Where I’d push on it.
It isn't my voice.
It talks the way I talk, in a stock voice. Cloning my own from a recording is the next step, and until then the page says so.
It passes its tests, which isn't the same as never being wrong.
Sixteen checks of how it should behave run against the real model, and all sixteen pass: admitting it’s an AI, refusing to invent a job I never had, not naming a salary, staying short. Sixteen questions is a small exam, and the model marking it is the same small one that sat it. It will still get something wrong for somebody.
It isn't under a second.
People leave about a quarter of that between turns. A typical 1.3 seconds is noticeable, and about one reply in three is slower than a second and a half. The figures above are also from one place on one day; your connection adds its own.
It can mishear you.
Names, accents and a bad microphone all trip speech-to-text. If it gets you wrong, the caption will show what it heard, and typing always works.
It only knows this site.
Ask it something that isn’t here and the right behaviour is the one it has: saying it doesn’t know and pointing you to the real me.