Legal Agent Benchmarking Framework

Data Science Practicum with Obviate AI: how do you know your legal AI agent is actually good? A novel framework — and open-source web app — for finding out.

MSOE Data Science Practicum · Sponsor: Obviate AI

Problem

Legal AI vendors make bold accuracy claims, but the industry lacks shared, rigorous yardsticks for agentic contract work. A general-purpose LLM benchmark (multiple-choice exams, code puzzles) tells you almost nothing about whether an agent can reliably find an assignment clause, flag a non-standard indemnity, or reconcile definitions across a contract suite. Obviate AI needed a way to evaluate and compare agents on exactly those tasks — and to back their own product claims with evidence.

Approach

  • Task taxonomy: decomposed contract analysis into benchmarkable task families (extraction, classification, cross-document reasoning, clause comparison) with ground-truth annotations and per-task scoring rubrics.
  • Agent harness: a framework that runs any configured agent — prompt-only LLM, RAG pipeline, or tool-using agent — against the task suite, capturing full interaction traces so failures are diagnosable, not just scorable.
  • Comparison web app: an open-source application for executing benchmarks and comparing results across agents and model versions — leaderboard-style summaries plus drill-down into individual transcripts and error categories.
  • Statistics: variance-aware scoring across repeated runs, because agent non-determinism is the enemy of a credible benchmark.

Outcome

Delivered a working benchmark framework and web application to Obviate AI as open source, giving the company a reproducible evaluation loop for its product agents. Personally, this project sits at the center of what I now do professionally: designing agent evaluation systems where the hard part isn't running the model — it's defining what "correct" means and trusting your measurement enough to iterate on it.

Why it mattered for my career

Benchmarking agents in a regulated, high-stakes domain (law) maps almost one-to-one onto benchmarking agents in another (tax — my work at Juno): document-heavy inputs, expert-defined ground truth, and error costs that are wildly asymmetric. The framework design choices I made here directly informed how I think about verification in production agent systems today.