The gap between a demo agent and a reliable one is usually not model intelligence; it is output discipline. In a demo, the model calls get_weather({"city": "SF"}) and looks magical. In production it calls refund_payment({"payment_id": 7712, "amount": "49.00", "currency": "usd", "reason": "user asked nicely"}) β€” integer where you expect a string, string where you expect cents, an undeclared reason, and an extra field your handler silently ignores. Every one of those mismatches is a bug report...