Problem
Legal AI vendors make bold accuracy claims, but the industry lacks shared, rigorous yardsticks for agentic contract work. A general-purpose LLM benchmark (multiple-choice exams, code puzzles) tells you almost nothing about whether an agent can reliably find an assignment clause, flag a non-standard indemnity, or reconcile definitions across a contract suite. Obviate AI needed a way to evaluate and compare agents on exactly those tasks — and to back their own product claims with evidence.
Approach
- Task taxonomy: decomposed contract analysis into benchmarkable task families (extraction, classification, cross-document reasoning, clause comparison) with ground-truth annotations and per-task scoring rubrics.
- Agent harness: a framework that runs any configured agent — prompt-only LLM, RAG pipeline, or tool-using agent — against the task suite, capturing full interaction traces so failures are diagnosable, not just scorable.
- Comparison web app: an open-source application for executing benchmarks and comparing results across agents and model versions — leaderboard-style summaries plus drill-down into individual transcripts and error categories.
- Statistics: variance-aware scoring across repeated runs, because agent non-determinism is the enemy of a credible benchmark.
Outcome
Delivered a working benchmark framework and web application to Obviate AI as open source, giving the company a reproducible evaluation loop for its product agents. Personally, this project sits at the center of what I now do professionally: designing agent evaluation systems where the hard part isn't running the model — it's defining what "correct" means and trusting your measurement enough to iterate on it.
Why it mattered for my career
Benchmarking agents in a regulated, high-stakes domain (law) maps almost one-to-one onto benchmarking agents in another (tax — my work at Juno): document-heavy inputs, expert-defined ground truth, and error costs that are wildly asymmetric. The framework design choices I made here directly informed how I think about verification in production agent systems today.