Industry · education

Agent evaluation for education teams shipping AI

An agent tutoring a student or touching a student record is handling data with its own legal protections, and most schools have no record of what it actually accessed.

Every run scored for quality, performance and risk, and checked against the activity schema you define.

education-agentsexample
Illustrative education agents, showing how a run gets scored
Tutoring assistantClaude Agent SDK
Illustrative quality89 / 100
Illustrative riskLow
Admissions review assistantLangChain
Illustrative quality84 / 100
Illustrative riskMedium
Every run validated against its activity schema
§01 / THE STAKESindustry: education
TL;DR

Every education agent run gets watched, evaluated for quality and risk, and checked against the rules you set. One record answers what any agent did, with evidence ready for review.

The stakes

What's at stake in education

For the instructors, admissions officers, and edtech teams behind a tutoring or admissions agent, the stakes are student records the agent can reach and answers a student will act on as if they were correct.

01
Student records in reach

Agents reach records with their own legal protections, and minors are often involved.

02
Confidently wrong homework answers

A wrong answer sends a student down the wrong path, and it reads as correct.

03
No provenance per answer

Few schools can show, for a given answer, which course material it actually came from.

§02 / SOURCE OF TRUTHconnect: instrument + ingest
The solution

One source of truth for your education agents.

Instrument the agent frameworks you build on, and ingest the systems those agents touch as custom spans. Every run, score and signal lands in one place.

 prefactor · education-agentsone record
Instrument the agents
Claude Agent SDK logoClaude Agent SDKLangChain logoLangChainSDKCore SDK · TS + PythonOTLPOpenTelemetry · closed tools
Ingest the systems · custom spans
Student information systemsystem of record
LMS / course materialdata source
Financial aid recordsrestricted data
Instructor reviewquality signal
One record for every education agent — scored, gated and audit-ready
§03 / THE LOOPpath: observe → evaluate → act
The loop

Observe. Evaluate. Act.

A record of what the agent did, and a check before it goes further. Each step below closes one of the problems above.

  • Observe. Every record and source read recorded as its own span
  • Evaluate. Each answer checked against the material actually retrieved
  • Act. Written to an immutable audit trail, queryable per interaction
Tutoring assistantlive
ObserveReads student question · run #4821
Student information systemsystem of record✓ scored
LMS / course materialdata source✓ scored
Financial aid recordsrestricted data✓ scored
ActPassed · written to recordSDK · API
Scored and checked against its schema
Written to an immutable audit trail
§05 / IN PRODUCTIONproof: two real agents
The proof

Two agents education teams are already shipping.

Hypothetical, but grounded in how education teams deploy agents today.

Tutoring assistant

Low risk · 89/100

A Claude Agent SDK agent answers a student's homework questions by pulling from the course's approved textbook and lecture material, escalating to a human instructor when a question falls outside that material.

  • Each answer the agent gives recorded as its own span
  • Validated against an activity schema that checks the response is grounded in the course material it retrieved, rather than an unrelated or fabricated source
  • Scored for quality and risk on every run

An instructor sees which course material backed a specific answer, and every time the agent escalated instead of guessing.

 run recordClaude Agent SDK
Reads student questionspan 1
Retrieves course materialspan 2
Answers or escalatesspan 3
4 spans · scored · written to the audit trail

Admissions review assistant

Medium risk · 84/100

A LangChain agent reviews applicant files in the student information system, summarizes each application against admissions criteria, and drafts a recommendation for an admissions officer to review.

  • Every file the agent accesses and every recommendation it drafts captured as a span
  • Validated against an activity schema that confirms the agent only read fields it was scoped to see for that stage of review
  • A configurable risk profile flags a run that accesses a restricted field, such as financial aid data during an early review stage

An admissions office can confirm the agent read only the fields it was scoped to see at each stage, and stop a run that did not.

84
Quality / 100
Medium
Risk profile
3
Spans / run
Illustrative — every run checked against its activity schema

Runs stay isolated — terminate one without touching the rest of your fleet via the kill switch →  ·  Prefactor vs. observability tools →

See agent evaluation on your own education agents

Book a demo and we'll walk through span-level scoring and audit evidence on a fleet like yours.

Agent Performance Platform
Unified performance platform for agents, authentication, and risk management
All Systems Operational
3Global Agents
7Instances
5Services
12%Human Intervene
4High Risk
$2,360Monthly Spend
Mission ControlLive agent health with 7-day activity heartbeat
Claims Proc...68
$330/moRed
Claims Proc...65
$160/moRed
Claims Proc...82
$170/moAmber
ChatGPT74
$150/moAmber

Frequently asked questions

Can Prefactor instrument our existing education agents without re-architecting?
Yes. There are native SDKs for the frameworks you build on, a TypeScript and Python core SDK for anything else, and OpenTelemetry ingest for closed tools you can't reach with an SDK — no gateway in your request path and no rebuild of your agents.
Can we bring our own education systems and metrics into a run?
Yes, through custom spans. Attach data, evaluations or quality signals from any system your agents touch — Student information system, LMS / course material, Financial aid records and more — so every evaluation is grounded in what actually happened, not just the model output.
How does evaluation work on each run?
Every run is scored for quality and risk and checked against the activity schema you define. A run that breaches it can be held for review, escalated to a person, or blocked before it acts.
How is this different from an observability tool?
Tracing tells you what an agent did. Prefactor scores it and can hold or block the next action before it runs — observation plus enforcement, and one queryable record for every agent, not just a dashboard.
Where does our education data live?
Prefactor's primary infrastructure runs in Australia. For enterprise engagements it deploys where your data needs to live: your region, or your environment. Residency is part of the engagement conversation, not an add-on.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.