The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
Why it earned a slot
The £85K chatbot PoC story is exactly the kind of war story the industry needs more of — six weeks just to build the evaluation and tracing infrastructure before the model was even selected. That's a brutal but honest indictment of how most enterprise AI projects are scoped. The stale embedding catch (interest rate policy update not reembedded → CSAT drop) is a perfect example of why observability isn't optional. The 47 PII breaches caught pre-launch alone makes this worth watching.
The short version
Your AI PoC isn't failing because of the model — it's failing because you skipped six weeks of infrastructure.
Why it matters
Enterprise agent deployment is shifting from 'can it work in a demo' to 'can it survive production and regulators.' The five-pillar framework Bhaumik lays out — evaluation, observability, data foundation, orchestration, governance — maps directly to what I see failing in most agentic systems I evaluate. The idea that the evaluation dataset is a living system, not a frozen benchmark, is the single most important shift practitioners need to internalize.
My take
Having built agentic systems end-to-end, I can tell you the evaluation-and-observability phase is where 80% of the real engineering happens. The model selection is the easy part. What Bhaumik's team did — instrumenting tracing before deployment, building a production incident playbook that ties all five pillars together — is what separates a production-grade agent from a science project. The governance piece (47 PII breaches caught before launch) is also the part most teams skip until a regulator forces them to face it.
How it connects
- The stale embedding failure mode is a specific instance of a broader problem: agentic systems need continuous data lineage tracking, not just point-in-time evaluation.
- European regulatory requirements for traceability (mentioned in the talk) are going to become a baseline expectation globally — this is early signal of what's coming.
- The production incident playbook concept bridges the gap between MLOps and agentops — teams building agents need the same rigor as teams running distributed systems at scale.
Bottom line
If your agent deployment doesn't have a tracing pipeline and a living evaluation dataset before you pick a model, you're not deploying — you're gambling.
Brendon Score: 8.5/10
- Quality: 8.2/10 — base
- Authority: 7.0/10 — +0.20
- Freshness: 1.3/10 — +0.00
- Engagement: 6.5/10 — +0.15
- Relevance: 9.0/10 — +0.00
- Sum: 8.55
- Total (published): 8.5/10
Why this is here
Checks cleared: relevance, slop-title-floor, authority (tier 7), embeddability.