← Back to blog

What Vercel and Bryo AI's Jev results show, and what they do not

What Vercel and Bryo AI's Jev results show, and what they do not
TL;DR

Vercel and Bryo AI report faster, cheaper results with Jev. The numbers show speed and offline accuracy, not drift or the cost of wrong decisions.

A benchmark win is not a production win

When TechCrunch reported that Vercel replaced OpenAI's Luna with Jev for command safety classification and saw latency drop by 5 to 18 times at p95 with greater accuracy, the reaction from founders and heads of AI was predictable: people started asking whether they should trial Jev. That question is reasonable. The reported numbers are real. But what a two-company early result can show you is much narrower than a swap decision requires, and understanding the gap between those two things is what this piece is for.

The gap itself has a shape. One April 2026 analysis found a 37 percent gap between lab benchmarks and real-world production performance for enterprise agentic systems. That is not a reason to dismiss early results; it is a reason to read them carefully before acting on them.

What the Vercel and Bryo AI results actually show

Vercel's result covers one task: command safety classification inside their `fx` tool. The improvement in latency, 5 to 18 times faster at p95, is meaningful for that specific use case because safety classification sits in the critical path of every command the tool handles. A faster classifier reduces the perceived response time of the whole system. The accuracy improvement is noted but not quantified in the reporting, so treat "greater accuracy" as a directional signal, not a figure you can carry into your own evaluation.

Bryo AI's CTO ran a more numerically detailed test: business email classification across 10 categories. Jev scored 96.4 percent accuracy. Gemini 3.5 scored 97.5 percent and Gemini 3.8 scored 98.5 percent. Jev was 10 to 20 times cheaper. That is genuinely useful information. If your task tolerates a 1 to 2 percentage point accuracy gap and your volume makes cost a real constraint, those numbers give you something to think with. If your task does not tolerate that gap, or if the gap behaves differently on your categories, the cost advantage does not automatically win.

Both results share three properties. First, they are offline evaluations: inputs were fixed, categories were known, and the model had no opportunity to interact with a live system in ways that would introduce distribution shift. Second, they cover narrow, well-defined tasks. Classification with a fixed label set is one of the most structured things you can ask a language model to do. Third, they reflect a single point in time. Neither result says anything about how performance changes as the model encounters edge cases, as your input distribution shifts, or as the task definition evolves.

These are not criticisms of Vercel or Bryo AI. Publishing early results is useful. The limitation is in what any offline benchmark can tell you, regardless of who runs it.

What those results cannot show

The three things the early Jev results leave unanswered matter more as your deployment gets closer to production.

Your input distribution is different. Vercel's command safety inputs and Bryo AI's business emails are not your data. Classification accuracy on 10 business email categories may look nothing like accuracy on your 10 categories, or your 4 categories, or your task where the label boundaries are less crisp. The gap between offline evaluation and production agent failures is where most swap decisions go wrong.

The cost of wrong decisions is not in the accuracy figure. A 96.4 percent accuracy rate means roughly 36 misclassifications per 1,000 inputs. What those 36 misclassifications cost depends entirely on what happens downstream. If a misclassified email gets routed to the wrong inbox, the cost is low. If a misclassified command passes a safety check it should have failed, the cost is not. Accuracy figures are averages. Silent agent failures in production are rarely distributed evenly across the label space; they tend to concentrate at the boundaries that matter most.

Drift is not visible in a point-in-time result. Neither result covers what happens over weeks or months as input patterns shift. Detecting agent quality decay and production drift requires continuous measurement against real traffic, not a one-time benchmark.

Speed gains interact with accuracy in ways that vary by architecture. Vercel's 5 to 18 times latency improvement is real, but the range (5x to 18x) is wide. P95 latency figures depend on your infrastructure, your call patterns, and where the classifier sits in your pipeline. Model routing and cost optimization in real agent workflows covers why headline latency figures often narrow in practice once you account for orchestration overhead.

This is not an argument against trialling Jev. It is an argument for knowing what you are measuring when you do.

What a useful comparison looks like

A benchmark run on someone else's data answers the question of whether Jev is worth testing. It does not answer the question of whether Jev works on your task, at your volume, with your error cost structure.

A useful comparison runs on real production traffic. You route a share of live inputs through both the current model and Jev, score the outputs against your actual quality criteria, and measure both accuracy and the downstream consequences of misses. That is different from running a held-out test set, because your held-out test set was drawn from historical data that may not reflect what users are sending today.

Offline evaluation versus production agent failures describes the mechanics of that transition in detail. The short version: fix the inputs, and you fix the hardest part of the problem for the model. Real inputs are not fixed.

When we built Prefactor's evaluation layer, the design decision we kept returning to was that scoring has to happen against live spans, not synthetic replays. The platform records every span your agent produces, scores quality and risk against criteria you define, and validates behaviour against activity schemas you set. That means when you run a model comparison, you are comparing performance on the inputs your system actually receives, including the ones that do not look like your training distribution.

One figure worth keeping in mind: benchmark success rates of 47 percent have been observed falling to 11 percent when agents operate on asynchronous real-world tasks. That is not a Jev-specific finding. It applies to any model swap evaluated primarily on offline benchmarks.

Agent benchmarks versus production, and the 37 percent quality gap covers the structural reasons for that drop in more detail. Evaluating agent readiness for autonomous deployment at enterprise scale gives a framework for deciding when a result is good enough to act on.

Where to start

Run your comparison on production traffic with your quality criteria, not on a held-out benchmark with Bryo AI's categories. Instrument your current setup first so you have a baseline, then add Jev to the same pipeline and score both against real inputs.

Start evaluating your agents with Prefactor's SDK to get span-level scoring on live runs, or read through the docs to see how activity schema validation and quality scoring work before you connect anything.

Frequently asked questions

Bryo AI's test showed Jev was 10 to 20 times cheaper than Gemini with only a small accuracy gap. Should I just switch?
The accuracy gap is 1 to 2 percentage points on Bryo AI's 10 email categories. Whether that gap is acceptable depends on what your misclassifications cost downstream, not on the percentage alone. Run the same comparison on your categories before treating the cost advantage as decisive.
Vercel reported 5 to 18 times faster latency. Why is the range so wide?
P95 latency figures vary with infrastructure, call patterns, and where the model sits in the pipeline. A 5x improvement and an 18x improvement describe different conditions, and the TechCrunch reporting does not specify which conditions produce which end of the range. Measure latency in your own pipeline before assuming you will land at the high end.
What is the minimum I need to set up before running a live model comparison?
You need a baseline: recorded outputs from your current model on real inputs, scored against your quality criteria. Without that baseline, you have no reference point to compare Jev against. Prefactor's SDK instruments your existing agent spans and scores them, which gives you that baseline before you introduce a second model.
How do I know if Jev's accuracy holds up over time rather than just at the point I tested it?
A one-time benchmark cannot tell you that. You need continuous scoring against live traffic, with alerting when accuracy or error patterns shift. That is what production drift monitoring covers, and it applies to any model, not just Jev.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.