Two Weeks at CallKaro: A Dropped Zero Broke a Voice Agent

· 3 min readPost views
Voice AIFDELLMCareerTesting

A week ago I wrote about what a Forward Deployed Engineer does. Two weeks in, I have a better answer: you find out, one phone call at a time, exactly how a voice agent fails. Here are the failures that taught me the most at CallKaro AI.

The best bug so far was a missing zero. A verification step asked callers for the last four digits of a number on file. One caller gave the right answer three times and got asked again every time. The value on file was 0408. The model passed it to the verification function as 408, the function saw three digits, treated that as a system error, and said "let me try that one more time." Nothing capped that path, so the loop never ended.

The fix was boring, which is how you know it was real. Pad the value back to four digits. Retry a system error once, then close the call politely instead of looping. Tell the function description to copy the value exactly, leading zeros included. Then replay the exact inputs from the failing call and check that it now passes. I did not trust the fix until the same call that broke it came back clean.

Another bug: after two failed attempts, the bot made up its own ending, "would you like to end the call now?" The correct scripted goodbye existed, but it lived in a separate step the model never moved to. So I stopped relying on the model to get there. The function now returns the scripted close itself on failure, and the prompt says to read it word for word. If a line has to be said exactly, make the tool hand it over.

A different kind of problem: the code was doing exactly what it was told, and what it was told was wrong. A voice-note briefing said one threshold, the written spec said another, and the implementation had drifted toward the voice note. I started keeping a contradiction sheet that lists each disagreement, who said what, and which source wins. The spec won every time, and the open questions went back to the business instead of being guessed at. The scariest bugs are the ones where the implementation is faithful to the wrong source.

Testing a voice agent is its own discipline. My setup is a pyramid. At the bottom are offline unit tests on the pre-call functions, which run with no network. Above them a wiring audit that greps for placeholders leaking into the prompt, then simulated conversations, then live test calls, then grep checks on the transcripts. Nothing moves up until the layer below is green, because a live call costs real money and real attention. One catch: the simulator cannot inject call metadata, so it only tests conversational rules. Anything scenario-specific needs a real call.

I also learned to be careful with numbers. After one prompt rewrite the base prompt came out 26% smaller and I nearly led with that. Then I measured the part that runs on most calls. It had doubled, so the worst-case path actually grew by about a third. We had moved tokens around, not removed them. That went into the write-up as a cost, not a win.

Smaller lessons that each ate an hour. On Windows, a Python script that prints Devanagari crashes the console unless you set PYTHONIOENCODING=utf-8. Updating an agent does nothing for live calls until you publish it as a separate step. And the standing rule I now follow without thinking: diagnose and propose first, never push to a live agent without an explicit go-ahead, and ask again before every test call.

Coming from embedded firmware, where a bug took days to reproduce, this loop is fast and a little scary. A bad prompt shows up in the next hour of calls. I would not call it easier, just louder. I'll write again once more of this has run in production for a while.