We use Jev, a judgment model from TypeSafe AI, as an advisory reviewer when merging work from parallel AI coding agents. Before each merge it answers a short batch of yes-or-no questions about the change. It can flag a merge for a closer human look, but it never approves, merges or pushes anything itself.
This post is about how we build Tempomat, not about the product. Jev is not part of Tempomat, and no customer data goes near it. It reads summaries of our own code changes. We're writing it up because "how do you trust work from many coding agents at once?" is a question a lot of teams have right now, and a small, cheap second opinion turned out to be a useful answer for us.
How do we build with parallel AI coding agents?
Our setup is simple to describe. One coordinating session acts as product manager: it splits the roadmap into written briefs and hands each brief to a worker agent. Each worker runs in its own git worktree, an isolated copy of the repository, so several can build at the same time without stepping on each other. When a worker is done, it commits on its own branch and reports back. The coordinator reviews the branch, runs the checks, and merges.
The bottleneck is the review. Every merge needs someone to ask: does this change do what the brief said, does it touch anything sensitive, and did it weaken a test to get green? Those questions are the same every time, which makes them a good fit for a model that answers questions instead of writing prose.
What is Jev?
Jev is a judgment model made by TypeSafe AI, which calls it "TypeSafe's first public System One Model, optimized for automation". It returns typed decisions with confidence scores. TypeSafe's coding-agent guide is direct about what it is not: "It does not generate text, write code, or hold a conversation." We reach it through the Experiential Labs AI gateway, the same gateway our main language model uses.
The main question type is a noul: a yes-or-no question for which Jev returns the probability that the answer is yes. A probability is easier to act on than a paragraph. You can compare it with a threshold in code, log it, and measure later how often it was right.
How does the merge-review gate work?
Before each worker merge, the coordinator makes one Jev call. The input is a compact state, not the whole diff: the diff statistics, the key facts about which files changed, the results of the checks, and the brief. Known-flaky test failures are listed separately from real failures, a lesson from the very first run.
The call asks a fixed set of questions, all in one batch:
- Merge or hold?
- Does the diff do what the brief asked?
- Does it touch authentication, row-level security, approvals or anything that can spend money?
- Does it weaken tests or evals?
- Does it change files outside the brief's scope?
- A risk score for the change as a whole.
The answers are printed and logged next to the coordinator's own verdict. Any disagreement means a closer human look, and the owner is told. The log keeps the question ids, the answers and the usage, plus a hash of the state instead of the state itself.
Batching matters for cost and speed. TypeSafe's parallel questions cookbook measured 13 questions in one call as 12.2× cheaper and 10× faster than asking them one by one. Its listed price, as of September 2026, is $0.042 per million input tokens, with output free. One of our live calls used 1,364 input tokens and took 1.2 seconds end to end, which works out to well under a hundredth of a cent.
What is the uncertainty band, and why use one?
A raw probability still needs a decision rule. We use a three-way band for every yes-or-no question:
| Probability of yes | What we do |
|---|---|
| Below 0.3 | Treat as no |
| 0.3 to 0.7 | Uncertain: a human looks |
| Above 0.7 | Treat as yes |
| High-stakes questions | Only act on yes above 0.85 |
The 0.3 and 0.7 cut-offs come from TypeSafe's self-consistency cookbook, which labels them illustrative. Its confidence routing pattern uses 0.85 as the line for acting alone on high-stakes actions. The point of the middle band is honesty: when the model isn't sure, the right output is "a person should check", not a coin flip dressed up as a verdict.
Two more habits come from the same reading. Arithmetic stays in code: TypeSafe's guide on building with System One models says to "keep deterministic work in code". And every question and threshold lives in one reviewed file, so changing what counts as "risky" is a code review of its own.
What can't the AI reviewer do?
The rules came from the owner before the first line of code, and they matter more than the model:
- It only adds checks or escalates to a human. It never approves, merges, pushes or executes anything by itself.
- One attempt, a bounded timeout, no retries. If Jev is slow or down, the merge goes ahead through the normal human review, as it did before Jev existed.
- Shadow mode first. For the first merges its answers are logged next to the human verdict and don't gate anything.
- Show the decision and the usage. Every answer and its token count is printed where the coordinator reads it.
Shadow mode is the part we'd recommend to anyone. Flavio Copes's write-up on Jev puts it plainly: run it beside your current process, log the answers, and label where it was right or wrong, with at least 20 realistic cases before you let it gate anything.
This mirrors a rule in the product itself. Tempomat never spends money without a person clicking Approve, and our reviewer never merges without a person deciding. Different stakes, same principle. If you're curious how we test the product's agent, see how we test an AI media buyer; for how we made it faster, our agent latency case study. What the product itself does is on the features page.
Where else do we use it? Context compaction
Long coordinator sessions run out of context and get compacted. We added a hook that asks Jev which of the session's notes are worth carrying over. In a dry run on a real coordinator session it looked at 58 candidates and kept 14, about 9,600 characters, using roughly 13,800 input tokens in 1.4 seconds. It decides what to remember; it doesn't decide what to do.
What did we decide not to use it for?
- Auto-approving read-only actions. It would have meant Jev granting permission, which breaks the first rule.
- Routing work to cheaper models. Every worker runs on the same model by the owner's choice.
- Our product's evals. Vercel's eve agent framework, which we build the product's agent on, uses Jev as its default eval judge through AI Gateway, according to eve's documentation. We don't have a Gateway key, so our evals use a small adapter over our own model instead.
- A third-party wrapper. We borrowed the tool design of an open-source Jev command-line tool rather than installing it, because it retries by default and targets other hosts.
Other teams use the same model for tighter loops. Firecrawl's explainer describes pi-warden, a guardrail for a coding agent, answering its checks "in about 250 ms" and holding 42 times over 17,000 recorded calls. Done-claim checks (did the agent say tests pass without running them?) and a risk gate for destructive shell commands are on our list, not built yet.
How will we know if it works?
After two weeks, or 20 merges, we'll count three things: real problems Jev flagged that the human review would have missed, false alarms, and how often Jev and the coordinator disagreed. Then each use stays, gets tuned, or goes. We'll publish what we find, including if the answer is "it wasn't worth it".
- 01Ask yes-or-no questions, not for a review essay.
- 02Batch every question into one call.
- 03Use a middle band that means "a human looks".
- 04Shadow mode first; label real cases before gating.
- 05The model escalates; people decide.
What is Jev?
Jev is a judgment model from TypeSafe AI. Instead of writing text, it answers typed questions, such as yes-or-no questions with a probability, which makes it suited to automated checks.
Can an AI review code merges?
It can help. We use it as an advisory second opinion: it answers fixed questions about each change and can flag a merge for a closer look, but a person makes the merge decision.
Does Jev approve or merge code by itself?
Not in our setup. It only adds checks or escalates to a human, with one attempt and no retries. If it's unavailable, the normal human review continues.
Is Jev part of the Tempomat product?
No. We use it only in how we build Tempomat. The product's own agent and its safety rules don't depend on it.
Tempomat is an AI media-buying copilot for Meta, TikTok and Google Ads, where nothing spends without your approval.
Start your 7-day free trial