What if the expensive AI is doing the wrong job?

We are testing whether a tiny decision model can handle the hundreds of semantic judgments surrounding frontier-model work — without giving it authority to act.

RUNNING LABAn experiment actively being exercised and measured. Results may change.
01 Real workload02 Deterministic pre-filter03 Jev judgment04 Confidence gate05 Existing model / human path06 Canonical outcome07 Compare and measure
01 · Question
Can TypeSafe AI's Jev reduce expensive model calls, context volume and human review by handling narrow classification, routing, scoring and verification decisions around real Lehvel workflows?
02 · Why it matters
Most AI systems spend frontier-model intelligence on two very different jobs at once: solving hard problems and making small semantic decisions around those problems. The second category is everywhere — is this evidence strong enough, is this result actually verified, does this source deserve synthesis, is this case routine or ambiguous? We want to know whether separating those jobs makes the whole system faster, cheaper and more reliable.
03 · State
RUNNING LAB An experiment actively being exercised and measured. Results may change.
04 · What we built or tested
  • A bounded Jev evaluation layer around existing Lehvel work rather than a new agent or orchestration system.
  • Jev through Vercel AI Gateway as typesafe-ai/jev, so the hosted path can use Vercel's existing identity boundary instead of spreading another provider key through the stack.
  • A division of labor: Jev judges, routes, scores and verifies; deterministic code enforces; frontier models create and reason.
  • Shadow/replay evaluation against real Kernel traces, execution reports and research workloads before any production shortcut is allowed.
  • Versioned question sets and explicit measurement of confidence, latency, cost, human corrections and high-confidence mistakes.
05 · Evidence
Provider shapestate in → typed Choice / Score / Boolean decisions plus probabilities; no prose generation required
Vercel pathtypesafe-ai/jev available through AI Gateway and AI SDK evaluate
Published input priceabout $0.04 per 1M input tokens through Vercel AI Gateway as of 2026-09-17
Lehvel resultnot measured yet — replay/shadow evaluation is starting now
06 · What we learned
  • The interesting opportunity is not replacing Claude or OpenAI. It is stopping frontier models from doing every small judgment surrounding the work.
  • Typed output is not truth. A valid high-confidence decision can still be wrong, so confidence never becomes permission.
  • The cleanest integration was through the runtime we already use. Putting Jev behind Vercel Gateway keeps it replaceable and avoids turning a model experiment into a new infrastructure stack.
07 · Limitations
  • TypeSafe and Vercel benchmark claims are vendor results, not Lehvel results. We will not repeat them as our performance.
  • No Lehvel workload has yet earned a production shortcut from Jev. Initial use is replay and shadow only.
  • Jev may advise on routing or verification, but it cannot choose a tenant or repository, grant permission, publish, contact a customer, mutate a public surface, move money or override deterministic policy.
  • If the evaluator adds another semantic hop without removing meaningful model cost, latency, rework or founder touches, we reject it.
08 · What changes next
Run Jev against real Lehvel Kernel, execution-verification and research cases; publish the measured agreement, high-confidence misses, latency, cost and frontier-model work actually avoided. Keep the parts that earn their place and delete the rest.
Negative results are published as results. A state on this page is a claim about what is true today, not about what is planned; a limitation is a limitation.

Tell us what is true about your business.

A few questions, in your own words — type them or talk. A person reads what you send and follows up. Nothing is published and nothing is sent on your behalf.