We need to talk about Jev
Matthew Berman's second video, eleven minutes, three days in. It plays where it was posted.
A browser agent that finds flights in seven seconds. A Claude session cut from a million tokens to 86,000 in one second. A WHERE clause that reads plain English. Vercel's safety reviewer. We are collecting what people built with Jev in its first days, and what they argued about, each in their own words. Nobody in these bylines is on eChai, and every entry links back to where it was posted.
8 entries Clear
Matthew Berman's second video, eleven minutes, three days in. It plays where it was posted.
We need to talk about Jev
Greg Isenberg, of Late Checkout, with the businessperson's explanation: Jev is the sorting part of a job, and a great many jobs are exactly that. The list of what it unlocks runs from instant quotes to inbound triage, and the episode of Startup Ideas with Ryan Vogel is linked. The film is 28 minutes and plays where it was posted.
Jev is HERE and this is the CLEAREST explanation of what it is and what NEW businesses it unlocks.
(and at the end I'll tell you how to get Jev even if you're on the waitlist)
WHAT IT IS
You know how you open your inbox and have to decide what's junk, what needs a reply, and what can wait? Jev does that part. It looks at each thing and says "this is junk, I'm 94% sure."
It doesn't write anything back to you. It just sorts.
1,700 emails for 18 cents, instantly.
That sounds kinda trivial but the important part
WHAT IT UNLOCKS
My explanation of Jev sounds small until you realize HOW MANY jobs are exactly this. Someone reading a stack of applications. Someone deciding which support ticket goes to which team. Someone looking at inbound and deciding who's worth calling back.
A few ideas on what it unlocks:
1/ Instant quotes that are actually instant. Every quote form on the internet says "we'll email you by end of day." Build the version that answers in under a second, for roofers, movers, insurance, legal intake.
2/ Lead scoring as a product. Every agency and service business has a contact form full of junk. Score every submission and send the real ones straight to the owner's phone.
3/ Support triage for companies with no support team. The ticket gets classified and routed before anyone opens it.
4/ Clipping tools. Pass in a transcript, get the best moments scored in three seconds. Every clipping product just got a cheaper engine.
5/ Application piles. Grants, permits, insurance claims, job apps, loan docs. Someone reads that stack one item at a time today.
6/ Marketplace matching. Someone types what they need and gets matched to the right local business instantly instead of waiting for callbacks.
7/ Browser agents that actually move FAST. That makes bulk browser work practical: pulling quotes from five carriers, filing the same form for 200 clients, checking supplier inventory in real time etc.
TLDR; find an expensive queue and put Jev at the front of it.
HOW TO GET IT
I didn't realize you can skip the waitlist because Jev is live on the Vercel AI Gateway right now, so you can start calling it today. In this episode, we share how.
Episode now live on @startupideaspod (thanks to @ryanvogel for coming on and spilling the sauce today)
Watch: https://www.youtube.com/watch?v=4mTLpuQpB80
Jev is a big deal because this is a whole new way to do AI
Really cool
Happy Jev day.
Matt Van Horn's summary of his own article, which checked every big Jev post of the first 72 hours by hand and groups them into nine things. Several entries on this page were found through it.
TL;DR of my new article: WTF is Jev by @typesafeai, and the 9 things people are already building with it. The thesis: 𝗮 𝗰𝗼-𝗰𝗿𝗲𝗮𝘁𝗼𝗿 𝗼𝗳 𝗖𝗵𝗮𝘁𝗚𝗣𝗧 𝘀𝗽𝗲𝗻𝘁 𝘁𝘄𝗼 𝘆𝗲𝗮𝗿𝘀 𝗶𝗻 𝘀𝘁𝗲𝗮𝗹𝘁𝗵 𝗼𝗻 𝗮 𝗺𝗼𝗱𝗲𝗹 𝘁𝗵𝗮𝘁 𝗰𝗮𝗻𝗻𝗼𝘁 𝘄𝗿𝗶𝘁𝗲 𝗮 𝘀𝗲𝗻𝘁𝗲𝗻𝗰𝗲, 𝗮𝗻𝗱 𝗶𝗻𝘀𝗶𝗱𝗲 𝟳𝟮 𝗵𝗼𝘂𝗿𝘀 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿𝘀 𝘄𝗶𝗿𝗲𝗱 𝗶𝘁 𝗶𝗻𝘁𝗼 𝗲𝘃𝗲𝗿𝘆 𝗰𝗵𝗲𝗮𝗽 𝗷𝘂𝗱𝗴𝗺𝗲𝗻𝘁 𝗰𝗮𝗹𝗹 𝗮𝗻 𝗮𝗴𝗲𝗻𝘁 𝗺𝗮𝗸𝗲𝘀.
Think AI multiple choice, not AI essay writing. It doesn't chat. You hand it app state plus a typed question, it hands back a decision with a probability attached. 𝟯𝟭.𝟰𝗠 𝘃𝗶𝗲𝘄𝘀 on the launch post in two days (@CompleteSkeptic, who co-invented RLHF). I ran @slashlast30days on it 11 times, then checked every big post by hand.
🌐 𝗔 𝘁𝗶𝗻𝘆 𝗼𝗽𝗲𝗻 𝘀𝗼𝘂𝗿𝗰𝗲 𝗯𝗿𝗼𝘄𝘀𝗲𝗿 𝗮𝗴𝗲𝗻𝘁 𝗳𝗼𝘂𝗻𝗱 𝗳𝗹𝗶𝗴𝗵𝘁𝘀 𝗶𝗻 𝟳 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 𝗳𝗼𝗿 $𝟬.𝟬𝟬𝟯𝟵. New action space every step, DOM as state, Jev picks the click, a small LLM only wakes up to type. The Browser Use founder built it (@gregpr07, 7.2K likes, 1.8M views) and had to note the video is 1x speed
🧹 The sleeper: instant compaction. Score every tool call, drop the junk, skip the summarization prompt entirely. "𝘪𝘯 2026, 𝘸𝘩𝘺 𝘪𝘴 𝘤𝘰𝘮𝘱𝘢𝘤𝘵𝘪𝘰𝘯 𝘴𝘵𝘪𝘭𝘭 𝘢 𝘴𝘶𝘮𝘮𝘢𝘳𝘪𝘻𝘢𝘵𝘪𝘰𝘯 𝘱𝘳𝘰𝘮𝘱𝘵?" asked @tamarajtran, 5K likes, then shipped the answer that afternoon. Run as a Claude plugin it took a session 𝗳𝗿𝗼𝗺 𝟭𝗠 𝘁𝗼𝗸𝗲𝗻𝘀 𝘁𝗼 𝟴𝟲𝗞 𝗶𝗻 𝗼𝗻𝗲 𝘀𝗲𝗰𝗼𝗻𝗱 (@altryne). Diogo's reply: "𝘧𝘳𝘦𝘦 𝘤𝘰𝘥𝘪𝘯𝘨 𝘢𝘨𝘦𝘯𝘵𝘴 𝘧𝘳𝘰𝘮 𝘥𝘦𝘴𝘪𝘨𝘯𝘪𝘯𝘨 𝘢𝘳𝘰𝘶𝘯𝘥 𝘵𝘩𝘦 𝘒𝘝 𝘤𝘢𝘤𝘩𝘦"
🛡️ Vercel put it in production as the safety reviewer in fx auto mode. 𝗨𝗽 𝘁𝗼 𝟭𝟴𝘅 𝗳𝗮𝘀𝘁𝗲𝗿 𝗮𝘁 𝗽𝟵𝟱 𝗮𝗻𝗱 𝗺𝗼𝗿𝗲 𝗮𝗰𝗰𝘂𝗿𝗮𝘁𝗲 than the model it replaced, per @rauchg, 3.7K likes. LangChain open-sourced the same idea the next day as AutoModeMiddleware. The closed danger classifier inside every coding harness is now a 100ms primitive
🚦 Model routing as middleware instead of a paragraph in a system prompt. About a dozen lines, probabilities left in agent state so you can audit the choice. The LangChain writeup by @sydneyrunkle is the cleanest how-to-wire-it piece anyone has published
🔎 RAG precision, solved the dumb way: retrieve as usual, run Jev on every chunk, delete the irrelevant ones. "𝘢𝘭𝘴𝘰 𝘥𝘪𝘥 𝘢𝘯𝘺𝘰𝘯𝘦 𝘳𝘦𝘢𝘭𝘪𝘻𝘦 𝘫𝘦𝘷 𝘴𝘰𝘭𝘷𝘦𝘥 𝘱𝘳𝘦𝘤𝘪𝘴𝘪𝘰𝘯 𝘪𝘯 𝘙𝘈𝘎?" (@kushbhuwalka, 416 likes)
🎮 Minecraft in real time: 𝗝𝗲𝘃 𝗿𝗲𝗮𝗰𝘁𝘀, 𝗚𝗣𝗧-𝟲 𝗔𝘀𝘁𝗿𝗮 𝗽𝗹𝗮𝗻𝘀, and they fight multiple zombies at once (@wuyang_zhou). A launcher that reads intent on every keystroke in about 100ms (@dabit3). TypeSafe's own demo is Doom at 10 decisions a second, roughly $7 an hour
📬 Email triage at scale: 1,500 emails in batches of 100 with 8 workers, 60,996 views on the demo. "𝘞𝘦 𝘰𝘯𝘭𝘺 𝘩𝘢𝘷𝘦 𝘢 𝘣𝘢𝘭𝘢𝘯𝘤𝘦 𝘰𝘧 $5 𝘥𝘰𝘸𝘯 𝘩𝘦𝘳𝘦, 𝘸𝘩𝘪𝘤𝘩 𝘫𝘶𝘴𝘵 𝘴𝘩𝘰𝘸𝘴 𝘩𝘰𝘸 𝘤𝘩𝘦𝘢𝘱 𝘵𝘩𝘪𝘴 𝘮𝘰𝘥𝘦𝘭 𝘪𝘴"
🗂️ 𝟳𝟳𝟳 𝗷𝘂𝗱𝗴𝗺𝗲𝗻𝘁𝘀 𝗶𝗻 𝘂𝗻𝗱𝗲𝗿 𝟬.𝟳 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 𝗳𝗼𝗿 𝗮 𝗾𝘂𝗮𝗿𝘁𝗲𝗿 𝗼𝗳 𝗮 𝗰𝗲𝗻𝘁. Every's head of evals asked 21 questions of 37 documents in one request, and that is what came back
🧪 Jev in your browser: Reflex, a Qwen model doing structured decisions on WebGPU, built at Shopify by @kshetrajna and passed around by @tobi. Three independent clones inside 72 hours. 𝗧𝗵𝗲 𝗶𝗻𝘁𝗲𝗿𝗳𝗮𝗰𝗲 𝗶𝘀 𝘁𝗵𝗲 𝗶𝗻𝘃𝗲𝗻𝘁𝗶𝗼𝗻, 𝗻𝗼𝘁 𝘁𝗵𝗲 𝘄𝗲𝗶𝗴𝗵𝘁𝘀
🔌 Already behind the gateways you use: @vercel AI Gateway inside 48 hours (2,341 likes, the company's second-biggest post), Cloudflare, and @OpenRouter in beta
💸 𝟱,𝟬𝟬𝟬 𝗿𝗲𝗾𝘂𝗲𝘀𝘁𝘀 𝗳𝗼𝗿 𝗮𝗯𝗼𝘂𝘁 $𝟮. That was one developer counting his bill on day one (@MichaelLee04, 3,060 likes). Input is $0.042 per million tokens. Output is free
🧨 The honest part: Every's second test came out 𝟮𝟱𝘅 𝗳𝗮𝘀𝘁𝗲𝗿, 𝗻𝗼𝘁 𝟮𝟬𝟬𝘅, and Jev caught 6 of 7 planted defects to Fable 5.1's 7. The HN launch thread (1,863 points) spent most of its length on "can't hallucinate." Top critical comment: "𝘪𝘵 𝘤𝘢𝘯'𝘵 𝘦𝘮𝘪𝘵 𝘢𝘯 𝘪𝘯𝘷𝘢𝘭𝘪𝘥 𝘵𝘺𝘱𝘦, 𝘣𝘶𝘵 𝘪𝘵 𝘤𝘢𝘯 𝘴𝘵𝘪𝘭𝘭 𝘦𝘮𝘪𝘵 𝘢 𝘤𝘰𝘮𝘱𝘭𝘦𝘵𝘦𝘭𝘺 𝘸𝘳𝘰𝘯𝘨 𝘷𝘢𝘭𝘪𝘥 𝘷𝘢𝘭𝘶𝘦." Diogo called the "it's a zero-shot classifier" read "𝘷𝘦𝘳𝘺 𝘢𝘤𝘤𝘶𝘳𝘢𝘵𝘦!" And the biggest Reddit thread is someone who open-sourced the same architecture a year ago, 1,568 upvotes. Top reply: "𝘉𝘶𝘵 𝘥𝘪𝘥 𝘺𝘰𝘶 𝘱𝘰𝘴𝘵 𝘪𝘵 𝘴𝘢𝘺𝘪𝘯𝘨 𝘪𝘵'𝘴 𝘵𝘩𝘦 𝘯𝘦𝘹𝘵 𝘣𝘪𝘨 𝘵𝘩𝘪𝘯𝘨? 𝘙𝘰𝘰𝘬𝘪𝘦 𝘮𝘪𝘴𝘵𝘢𝘬𝘦"
Bonus: the name is not Kahneman. It's William Stanley Jevons, of Jevons paradox. Make a resource cheaper and people consume far more of it. Naming your decision model after that is a thesis statement.
𝗞𝗲𝗲𝗽 𝘁𝗵𝗲 𝗯𝗶𝗴 𝗺𝗼𝗱𝗲𝗹 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗵𝗮𝗿𝗱 𝘁𝗵𝗶𝗻𝗸𝗶𝗻𝗴 𝗮𝗻𝗱 𝘄𝗿𝗶𝘁𝗶𝗻𝗴. 𝗨𝘀𝗲 𝗝𝗲𝘃 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗿𝗮𝗽𝗶𝗱-𝗳𝗶𝗿𝗲 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻𝘀 𝗶𝗻 𝗯𝗲𝘁𝘄𝗲𝗲𝗻. That's the whole article.
Moritz Kremb's 22-minute tutorial: what Jev is, the API, then three demos, a voice-controlled browser, an AI memory and a YouTube predictor. It plays where it was posted; the voice browser is open source.
Full Jev Tutorial
What it is, how you can build with it and what new applications it can unlock
→ 0:00 Intro
→ 0:34 Jev explained
→ 4:06 API setup
→ 5:59 Demo 1: Voice-controlled browser
→ 11:33 Demo 2: AI memory
→ 17:27 Demo 3: YouTube predictor
David Fant, on how an agent gets ten times faster and cheaper with Jev in it: routing, computer use, action review, and subagent orchestration, each pointing at somebody's demo.
jev will make agents 10x faster and cheaper, here's how:
1/ model routing: pick the right model for each task, without training a custom router
https://x.com/mdlahfir/status/2100314182201802811?s=20
2/ computer use: faster, cheaper and more reliable for action-heavy tasks
https://x.com/gregpr07/status/2100411066966749359
3/ auto review: ask jev whether an action is safe, instead of using a slow and expensive LLM
https://x.com/fazxes/status/2100300097695232164?s=20
4/ less obvious: subagent orchestration
long-running agents (cursor projects, grokbot, energy) parallelize work with subagents.
but every user message, email, or subagent reply can wake the expensive orchestrator.
example: it costs $1 to wake up gpt 6 astra w 100k input tokens
jev can decide what each event needs:
- route directly to a subagent
- queue for later
- wake the orchestrator
Paarangat's is the shortest version: a language model reasons its way to a sentence, Jev fills in decisions you defined in advance, each with a probability.
this is the easiest way to understand Jev:
LLMs generate answers.
Jev makes decisions.
that sounds like a small difference, but it actually changes the entire use case.
say you give a normal LLM this:
“here’s a user, their account history, payment behavior, support chats, device data, etc.
tell me if this looks risky.”
the LLM might reason through it and return:
“yes, this looks high risk.”
maybe in JSON if you ask nicely.
with Jev, you define the possible decisions upfront:
risk:
* low
* medium
* high
manual review:
* yes
* no
and Jev returns something closer to:
risk = high (96%)
manual review = yes (91%)
that’s basically the product.
it’s not trying to be another ChatGPT.
it’s more like an AI-native if statement.
instead of:
if transaction > $10,000:
review()
you can start thinking more like:
if “does this behavior look suspicious?” > 95%:
review()
and that opens up a pretty interesting category of software.
a few assumptions I had at first that turned out to be wrong:
1. “so it’s just a classifier?”
kind of, but that undersells it.
the input can be messy real-world context, and you can ask multiple typed questions about that state at once.
fraud?
churn?
escalate?
eligible?
priority?
all from the same input.
2. “so it replaces GPT / Claude?”
not really.
I actually think the interesting architecture is:
Jev decides WHAT needs to happen
Claude / GPT reason or generate WHEN deeper intelligence is needed
normal code executes the deterministic stuff.
Jev becomes the routing layer.
3. “it can’t hallucinate?”
this one needs nuance.
if your allowed answers are:
LOW
MEDIUM
HIGH
Jev won’t suddenly invent:
“EXTREMELY HIGH 🚨”
the output structure is constrained.
but it can still be wrong.
HIGH at 92% can still be the wrong decision.
so “no hallucinations” doesn’t mean “always correct.”
4. “why not just force an LLM to return JSON?”
you can.
we already do this everywhere.
but you still deal with generation latency, schema validation, retries, weird outputs, confidence estimation and a lot of glue code.
Jev is designed around the decision itself rather than text generation.
5. “why should I care?”
because most software is ultimately a giant tree of:
if this → do that
if this → route here
if this → escalate
if this → reject
if this → ask a human
Jev is basically asking:
what if those if statements could understand messy human context?
that’s a much more interesting framing than “another AI model.”
I can see this being very useful for:
fraud / risk
support routing
moderation
PR / QA automation
lead scoring
compliance
workflow orchestration
agent routing
especially as the cheap + fast decision layer sitting in front of larger reasoning models.
early tech, obviously.
but the category itself makes a lot of sense.
Nathan Flurry is a co-founder of Rivet. This is the explanation most people ended up passing around, and it is worth reading whole: what Jev cannot do, what it can, and the shape of workflow it fits.
hype-free explanation of jev:
jev does not replace gpt / claude
jev is just a *really* smart switch statement
like if 2016 ml classifiers got 2026 levels of intelligence
it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate
* = and by new, i mean rebranded
~~~
it needs a predefined set of options and it will tell you which one to take
it cannot:
- write code
- generate natural language
- reason step by step / show its work
- produce any output you didn't define in advance
- pick from more than ~255 options in one shot
but it can:
- classify, route, score, rank
- give confidence
- pick the right branch, tool, model, or sub-agent
- judge / verify / guardrail an llm's output
- label tons and tons of rows
~~~
i'd imagine a lot of workflows that look like:
llm proposes options → jev decides → code executes
and i see this fitting *really* well with code mode and mcp
~~~
implying this will lead to agi seems incredibly far fetched to me, but i don't want to discount the types of applications that this will make possible
Michael Lee builds a chat product and had a personal benchmark waiting, having spent months on the 'should we send a follow-up' class of question with DeepSeek Flash and GPT 5.6 Luna. This is the most complete first-hand account of using it in a product on the page: 5,000 requests for about $2, p50 around 150 ms, and a run of design notes, including that Jev rewards splitting a query into independent questions where an LLM punishes it.
I got access to Jev earlier today (thank you @hackgoofer). I have run ~5,000 requests so far, (which cost me around $2!), across classification, model routing, intent, steering, and many other things.
tl;dr, Jev enables a new intelligent decision-making primitive, separate from deterministic code and LLM calls. This allows a class of decision-making that was neither suited to dumb, unintelligent code, nor to slow, expensive LLMs.
It is super fast and cheap, and I think I will likely end up making a few Jev calls to every LLM call I make in my product. I think probably any company using LLM requests today can probably add a Jev call pre and/or post LLM calls to quite literally make their product much better for free, and have better tool calling behavior in many cases.
I happened to have a personal benchmark for this as I’d been working on a ton of proactivity and classification tasks. I have been using the deepseek flash and more recently gpt 5.6 luna family of models as reasonably smart classifiers with low latency. Think questions like:
- Did this conversation output contradict something they’ve mentioned before?
- Should we send a followup message to this user based on our rules?
- It’s been a few seconds of silence. Should we proactively send a message?
In the past, I’ve been forced to write a bunch of what I call decision chains, mostly because an LLM classification call is very expensive in TIME (avg 4s), and less importantly can cost quite a bit if run on every message.
Imagine a normal chat app. If you added 4s to every response to figure out if the response is good before sending it out, that ends up being pretty bad. So instead, I usually have to write some code that is a crude heuristic that runs quickly and decides whether to run the classifier. Obviously this sucks because you call the classifier many times that you don’t want to, which makes your p95 bad, and you also miss cases with the heuristic, and you also have to manage all of these weird chains.
With Jev, it’s cheap enough, and fast enough (p50 ~150ms, p95 ~350ms in my testing!) that you can easily run it every turn. Heck you can reasonably run it before generation AND post generation, for any application that isn’t realtime voice, and still feel snappy.
But this is just one use case. Think: smarter model routing, better context packing, better responses, better observability for intent/tags/safety, smarter retries and so much more. By simply thinking about the inputs and outcomes you want to enforce, you can use Jev to supercharge most model calls and reduce bad user outcomes. The more “quirks” a model has, the more valuable it ends up being.
It’s a bit weird and unintuitive using Jev. Generally, you want to decrease the # of questions you ask a classifier, or it makes more mistakes. In fact, you might want to ask your questions kind of in a compound way, because the reasoning happens in a shared scratchpad of sorts. Adding questions muddies the scratchpad and makes it take longer.
With Jev, you feel incentivized to go the other way, to formulate your query as a set of independent questions. It doesn’t feel like adding more questions decreases your performance on others.
You can go a bit deeper to improve tool calls. Many model tools are things like turning on settings, or other things. You can easily improve models that are not very good at tool calling with Jev, by simply figuring out when to run them. You can do a pre-LLM call to figure out when to unfurl different tool definitions, in order to make your main LLM run better, you could run a background task with Jev + another LLM to reduce tool and context burden on your main LLM, and free it to be responsive.
I’ve only scratched the surface of my testing but very excited!