Industry · legal

Agent evaluation for legal teams shipping AI

An agent that drafts a memo or reviews a contract is touching privileged information and citing sources a partner will be held to, and most firms have no record of how it got there.

Every run scored for quality, performance and risk, and checked against the activity schema you define.

legal-agentsexample
Illustrative legal agents, showing how a run gets scored
Contract review agentLlamaIndex
Illustrative quality91 / 100
Illustrative riskLow
Legal research memo agentClaude Agent SDK
Illustrative quality86 / 100
Illustrative riskMedium
Every run validated against its activity schema
§01 / THE STAKESindustry: legal
TL;DR

Every legal agent run gets watched, evaluated for quality and risk, and checked against the rules you set. One record answers what any agent did, with evidence ready for review.

The stakes

What's at stake in legal

For the partners, associates, and engineering teams behind a drafting or research agent, the stakes are privileged client information the agent can reach and citations a partner is professionally responsible for.

01
Privileged data in reach

Agents reach documents under privilege, and the firm needs a record of what each one read.

02
Plausible but wrong citations

A fabricated or misapplied citation is real professional exposure, and it looks plausible on the page.

03
No provenance per memo

Few firms can show, for a given memo, where each citation actually came from.

§02 / SOURCE OF TRUTHconnect: instrument + ingest
The solution

One source of truth for your legal agents.

Instrument the agent frameworks you build on, and ingest the systems those agents touch as custom spans. Every run, score and signal lands in one place.

 prefactor · legal-agentsone record
Instrument the agents
LlamaIndex logoLlamaIndexClaude Agent SDK logoClaude Agent SDKSDKCore SDK · TS + PythonOTLPOpenTelemetry · closed tools
Ingest the systems · custom spans
Case-law databasedata source
Document managementsystem of record
Contract playbookpolicy
Associate reviewquality signal
One record for every legal agent — scored, gated and audit-ready
§03 / THE LOOPpath: observe → evaluate → act
The loop

Observe. Evaluate. Act.

A record of what the agent did, and a check before it goes further. Each step below closes one of the problems above.

  • Observe. Every document read recorded as its own span
  • Evaluate. Each citation checked against the case actually retrieved
  • Act. Written to an immutable audit trail, queryable per matter
Contract review agentlive
ObserveReads contract document · run #4821
Case-law databasedata source✓ scored
Document managementsystem of record✓ scored
Contract playbookpolicy✓ scored
ActPassed · written to recordSDK · API
Scored and checked against its schema
Written to an immutable audit trail
§05 / IN PRODUCTIONproof: two real agents
The proof

Two agents legal teams are already shipping.

Hypothetical, but grounded in how legal teams deploy agents today.

Contract review agent

Low risk · 91/100

A LlamaIndex agent reviews incoming NDAs and vendor contracts against a firm playbook, extracts key clauses, and flags deviations from standard terms for an associate to review before redlining.

  • Each clause extraction and each flagged deviation recorded as its own span
  • Validated against an activity schema that checks the extracted clause actually appears in the source document, rather than being paraphrased or invented
  • Scored for quality and risk on every run

For any flagged deviation, an associate sees exactly which clause and document it came from, before it reaches a client.

 run recordLlamaIndex
Reads contract documentspan 1
Extracts key clausesspan 2
Flags playbook deviationsspan 3
4 spans · scored · written to the audit trail

Legal research memo agent

Medium risk · 86/100

A Claude Agent SDK agent researches a legal question across a case law database, drafts a memo with citations, and routes it to an associate for review before it reaches a partner.

  • Every case retrieved and every citation written into the memo captured as a span
  • Each citation checked against its activity schema: whether the case actually exists, and whether the retrieved text supports the point the memo attributes to it
  • A configurable risk profile flags memos with a high rate of unverified citations for closer review

The associate sees which citations were verified against the source case and which were not, before it goes to a partner.

86
Quality / 100
Medium
Risk profile
3
Spans / run
Illustrative — every run checked against its activity schema

Runs stay isolated — terminate one without touching the rest of your fleet via the kill switch →  ·  Prefactor vs. observability tools →

See agent evaluation on your own legal agents

Book a demo and we'll walk through span-level scoring and audit evidence on a fleet like yours.

Agent Performance Platform
Unified performance platform for agents, authentication, and risk management
All Systems Operational
3Global Agents
7Instances
5Services
12%Human Intervene
4High Risk
$2,360Monthly Spend
Mission ControlLive agent health with 7-day activity heartbeat
Claims Proc...68
$330/moRed
Claims Proc...65
$160/moRed
Claims Proc...82
$170/moAmber
ChatGPT74
$150/moAmber

Frequently asked questions

Can Prefactor instrument our existing legal agents without re-architecting?
Yes. There are native SDKs for the frameworks you build on, a TypeScript and Python core SDK for anything else, and OpenTelemetry ingest for closed tools you can't reach with an SDK — no gateway in your request path and no rebuild of your agents.
Can we bring our own legal systems and metrics into a run?
Yes, through custom spans. Attach data, evaluations or quality signals from any system your agents touch — Case-law database, Document management, Contract playbook and more — so every evaluation is grounded in what actually happened, not just the model output.
How does evaluation work on each run?
Every run is scored for quality and risk and checked against the activity schema you define. A run that breaches it can be held for review, escalated to a person, or blocked before it acts.
How is this different from an observability tool?
Tracing tells you what an agent did. Prefactor scores it and can hold or block the next action before it runs — observation plus enforcement, and one queryable record for every agent, not just a dashboard.
Where does our legal data live?
Prefactor's primary infrastructure runs in Australia. For enterprise engagements it deploys where your data needs to live: your region, or your environment. Residency is part of the engagement conversation, not an add-on.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.