Some lessons only land when you have to retire something you were proud of.
What we shipped: a fully autonomous agent that could approve low-risk customer service refunds without human review. Cost-justified, well-scoped, passed every test.
Genuine question because the answer keeps changing.
Six months ago GPT-4o was the default for most teams. Then Claude 3.5 Sonnet started winning reasoning-heavy use cases. Now Llama 3 is showing up in production for cost-sensitive or on-prem deployments.
Sharing because these patterns are too consistent to ignore.
Mistake 1: Picking a model before understanding the use case
Teams pick GPT-4o because it's the default. Then realise their workflow needs structured output that Claude handles better, or cost constraints that only Llama satisfies. Model choice should come last, not first.