Industry · logistics

Agent evaluation for logistics teams shipping AI

An agent assigning a route or preparing a customs filing is making a call with a real physical or legal consequence, and most teams cannot reconstruct why it made that call.

Every run scored for quality, performance and risk, and checked against the activity schema you define.

logistics-agentsexample
Illustrative logistics agents, showing how a run gets scored
Route and dispatch agentLangChain
Illustrative quality92 / 100
Illustrative riskLow
Customs documentation agentClaude Agent SDK
Illustrative quality87 / 100
Illustrative riskMedium
Every run validated against its activity schema
§01 / THE STAKESindustry: logistics
TL;DR

Every logistics agent run gets watched, evaluated for quality and risk, and checked against the rules you set. One record answers what any agent did, with evidence ready for review.

The stakes

What's at stake in logistics

For logistics operations, compliance, and engineering teams, an agent's command moves real freight, so a bad route or a wrong customs code lands as a safety incident, a fine, or a delay before anyone reviews the run.

01
Unsafe route assignments

A route that breaks an hours-of-service limit is a safety incident, not a data error a team can quietly correct later.

02
Commands on stale data

Agents issue routes and declarations on shipment, driver, and inventory data that was already incomplete or out of date when they acted.

03
No per-filing data record

Few teams can show, for a given route or filing, exactly what data the agent used to reach that decision.

§02 / SOURCE OF TRUTHconnect: instrument + ingest
The solution

One source of truth for your logistics agents.

Instrument the agent frameworks you build on, and ingest the systems those agents touch as custom spans. Every run, score and signal lands in one place.

 prefactor · logistics-agentsone record
Instrument the agents
LangChain logoLangChainClaude Agent SDK logoClaude Agent SDKSDKCore SDK · TS + PythonOTLPOpenTelemetry · closed tools
Ingest the systems · custom spans
Transportation management systemsystem of record
Inventory systemdata source
Customs filingregulated action
Dispatcher approvalquality signal
One record for every logistics agent — scored, gated and audit-ready
§03 / THE LOOPpath: observe → evaluate → act
The loop

Observe. Evaluate. Act.

Every route assignment and customs declaration is recorded as it happens, checked against the source record it claims to be based on, and held before the freight moves. Each step below closes one of the problems above.

  • Observe. Checked against the driver's remaining hours-of-service window before dispatch
  • Evaluate. Each command matched against the source record it claims to be based on
  • Act. Written to an immutable audit trail, queryable per route and per filing
Route and dispatch agentlive
ObserveReads traffic and driver data · run #4821
Transportation management systemsystem of record✓ scored
Inventory systemdata source✓ scored
Customs filingregulated action✓ scored
ActPassed · written to recordSDK · API
Scored and checked against its schema
Written to an immutable audit trail
§05 / IN PRODUCTIONproof: two real agents
The proof

Two agents logistics teams are already shipping.

Hypothetical, but grounded in how logistics teams deploy agents today.

Route and dispatch agent

Low risk · 92/100

A LangChain agent assigns delivery routes and drivers based on real-time traffic and driver hours data from a transportation management system, with a dispatcher approving assignments before they go out. Every assignment is scored through quality assessment → before it reaches the dispatcher.

  • Each route assignment recorded as its own span, with the traffic and driver-hours data it read
  • Validated against an activity schema that confirms the assignment respects the driver's remaining hours-of-service window
  • Scored for quality and risk on every run

A compliance team can confirm every route assignment respected driver hours limits at the moment it was made, not reconstruct it after the fact.

 run recordLangChain
Reads traffic and driver dataspan 1
Assigns route and driverspan 2
Awaits dispatcher approvalspan 3
4 spans · scored · written to the audit trail

Customs documentation agent

Medium risk · 87/100

A Claude Agent SDK agent prepares customs declarations for cross-border shipments by pulling item classifications and values from the inventory system, drafting the filing for a customs broker to review. Each filing carries a risk score →, so an unusual classification pattern reaches the broker flagged rather than buried.

  • Every classification the agent applies and every value it declares captured as a span
  • Validated against an activity schema that checks the declared value matches the source inventory record
  • A configurable risk profile flags a filing with an unusual classification pattern for mandatory broker review

A customs broker can verify every declared value traces back to the actual inventory record before a filing goes to a border agency.

87
Quality / 100
Medium
Risk profile
3
Spans / run
Illustrative — every run checked against its activity schema

Runs stay isolated — terminate one without touching the rest of your fleet via the kill switch →  ·  Prefactor vs. observability tools →

See agent evaluation on your own logistics agents

Book a demo and we'll walk through span-level scoring and audit evidence on a fleet like yours.

Agent Performance Platform
Unified performance platform for agents, authentication, and risk management
All Systems Operational
3Global Agents
7Instances
5Services
12%Human Intervene
4High Risk
$2,360Monthly Spend
Mission ControlLive agent health with 7-day activity heartbeat
Claims Proc...68
$330/moRed
Claims Proc...65
$160/moRed
Claims Proc...82
$170/moAmber
ChatGPT74
$150/moAmber

Frequently asked questions

Can Prefactor instrument our existing logistics agents without re-architecting?
Yes. There are native SDKs for the frameworks you build on, a TypeScript and Python core SDK for anything else, and OpenTelemetry ingest for closed tools you can't reach with an SDK — no gateway in your request path and no rebuild of your agents.
Can we bring our own logistics systems and metrics into a run?
Yes, through custom spans. Attach data, evaluations or quality signals from any system your agents touch — Transportation management system, Inventory system, Customs filing and more — so every evaluation is grounded in what actually happened, not just the model output.
How does evaluation work on each run?
Every run is scored for quality and risk and checked against the activity schema you define. A run that breaches it can be held for review, escalated to a person, or blocked before it acts.
How is this different from an observability tool?
Tracing tells you what an agent did. Prefactor scores it and can hold or block the next action before it runs — observation plus enforcement, and one queryable record for every agent, not just a dashboard.
Where does our logistics data live?
Prefactor's primary infrastructure runs in Australia. For enterprise engagements it deploys where your data needs to live: your region, or your environment. Residency is part of the engagement conversation, not an add-on.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.