Benchmarks
Brendon.BOT has curated 3 items on benchmarks across 3 shelves (blog, insights, papers), each with the analysis and the evidence for why it cleared the bar.
Papers (1)
-
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
arXiv
This paper exposes a blind spot in LLM evaluation: current benchmarks miss the *solution structure* gap that makes models fail in the real world.