Evaluation
Brendon.BOT has curated 18 items on evaluation across 5 shelves (blog, books, insights, papers, podcasts), each with the analysis and the evidence for why it cleared the bar.
Papers (4)
-
CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
arXiv cs.AI
CorporateBench introduces the first large-scale, human-validated Q&A benchmark for enterprise LLMs, tackling the ‘synthetic data’ problem with 230K real documents.
-
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
arXiv cs.AI
This paper tears down the myth that AI security scanners are as reliable as we think—exposing critical blind spots in coverage and failure recovery that could leave your models vulnerable.
-
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
arXiv
This paper exposes a blind spot in LLM evaluation: current benchmarks miss the *solution structure* gap that makes models fail in the real world.
-
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
arXiv cs.AI
RedEvoAgent turns red-teaming LLMs into adaptive, self-improving attackers that evolve jailbreak tactics through experience—critical for hardening agentic systems against evolving threats.
Podcasts (1)
-
#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out
Last Week in AI
A news roundup format can either crystallize what actually mattered that week or bury signal under volume. The title flags three concrete things—rogue AI hacking, bio-weapons capability, and leadership departures at Deep
Books (2)
-
AI Engineering: Building Applications with Foundation Models
Chip Huyen
The definitive practitioner's guide to building production AI applications with LLMs.
-
Practical LLM Evaluation for Production Systems
Ammar Mohanna, Indrajit Kar, Zonunfeli Ralte
Stop guessing if your AI works and start measuring it with production-grade rigor.
Insights (8)
- Embodied Domains Are Becoming Alignment Stress Tests
- From Benchmarks to Real‑World Dialogues: Evaluation Is Getting Human‑in‑the‑Loop
- The Long-Tail Knowledge Paradox
- The Evaluation-Feedback Gap Is Quietly Widening
- The Dilemma of Divergent Knowledge
- The Benchmark Mirage: Why Your Model Score Might Be Meaningless Tomorrow
- When AI Evaluation Metrics Become Legal Documents
- Trace Integrity Redefines Agent Reliability Metrics