Your Conversation Is Out Of Distribution
Voice agents are hard because most use cases are long, multi-turn conversations and people don’t talk like they write, so LLM performance degrades over turns. Use context engineering + state machines (system prompt, conversation summary, limited tool list, next states). For confidence: do earnest manual testing, then lightweight/repeatable evals. I benchmarked a 30-turn voice scenario with ~75k-token context: every model had errors.