This is a second starter post. Replace it with your own content. It demonstrates linking back to related posts and tagging.
The demo-to-production gap
A tool call that works 95% of the time in a notebook is a different product than one that needs to work at 99.9% in front of customers. The last few percentage points usually come from handling cases nobody wrote a test for.
Where it breaks
- Ambiguous tool schemas: the model guesses at a parameter the schema didn’t constrain tightly enough.
- Silent partial failures: a tool call “succeeds” but returns something the model wasn’t expecting, and it proceeds anyway.
- No retry semantics: transient failures get treated as terminal ones.
- Context rot: long-running agent loops accumulate stale or contradictory context that biases later tool calls.
Practical fixes
- Make invalid states unrepresentable in the tool schema, not just documented as constraints.
- Return structured errors the model can reason about, not raw stack traces.
- Log every tool call and result. You cannot debug an agent you cannot replay.
- Set explicit budgets (steps, cost, time) so failures degrade gracefully instead of looping forever.
See also: Agent Orchestration Patterns