Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
This is no longer on the current shelf — shelves rotate as new material clears the bar. The analysis below is unchanged. See what is featured now.
Why it earned a slot
The 'vibes problem' framing hits hard — most teams shipping agents are literally just running queries and hoping for the best. That 0/13 versus 13/13 contrast between correctness and faithfulness evals is the kind of thing that should make every ML engineer re-examine their own eval setup. Laurie walks through the whole pipeline from Phoenix tracing to custom rubrics, and the real gem is the experimentation section most eval content skips entirely.
The short version
Your agent passes your eval. It's still lying to your users.
Why it matters
Agent evals are the single biggest unsolved bottleneck in production agentic systems right now. Most teams either don't have evals at all or rely on crude correctness checks that miss exactly the failures that matter — like a model confidently generating fabricated financial data it can't verify. This workshop directly addresses that gap with a practical, full-stack approach.
My take
The eval strategy you choose at the start dictates everything downstream — architecture, prompt design, even how you version your models. The 0/13 versus 13/13 finding Laurie highlights is exactly the trap teams fall into: they optimize for the wrong signal and ship something that looks great in testing but fails silently in production. The emphasis on root-cause categorization before writing any eval code is the discipline most teams skip because it's slower and less glamorous.
How it connects
- As LLM-as-judge evals become standard, the real differentiator is knowing when NOT to trust the judge — and this workshop addresses that tension head-on
- The experimentation framework for proving prompt changes actually worked closes the loop between research and production that most agent teams leave open
- Financial domain agents are an especially punishing test case because ground truth is often unverifiable, making eval design existential rather than optional
Bottom line
Build your eval pipeline before you build your agent — and make sure your eval actually measures what you think it measures.
Brendon Score: 8.3/10
- Quality: 8.0/10 — base
- Authority: 7.0/10 — +0.20
- Freshness: 1.0/10 — +0.00
- Engagement: 6.4/10 — +0.14
- Relevance: 9.0/10 — +0.00
- Sum: 8.34
- Total (rounded): 8.3/10
Why this is here
Checks cleared: relevance, slop-title-floor, authority (tier 7), embeddability.