What happens when enterprise requirements - human approval gates, audit trails, structured output - hit three agent frameworks? The first article measured how Strands, LangGraph, and CrewAI differ on a plain task. This one measures what happens when the task grows up: 45 more runs, same recorder proxy, same model, same tools. The headline finding: the frameworks fail differently. Strands - the model-driven one - finished with an empty output three times while exiting LangGraph and...