Grep, Then Jev: A $0.04 Judge for Picking the Right Code Files
In We Measured Grep on Real Codebases we found that grep rarely misses the right files. It buries them. The step that decides which grep hits matter is where the quality is won or lost, and it’s also where the money goes: today that step is usually a frontier model reading big batches of code.
At the end of that post we listed “fast classifier models as the judge” as untested. This is that test.
We put TypeSafe’s Jev, a small decision model, behind grep. We ran it over 5,075 candidate files from six codebases and compared its picks with Opus 5.5’s on exactly the same files.
The short answer: Jev found more of the right files than Opus at every depth. With a stricter threshold it matched Opus’s precision. It cost about 1% as much.
The problem
Ask a coding agent “How are taxes calculated?” in Saleor and it will grep for tax. Grep returns 107 files. Only 19 of them are files a developer actually changes when tax logic changes. The rest import zero_money from core.taxes, log a tax error, or name a field tax_rate.
Someone has to read those 107 files and decide which ones matter. That judgment step is what we measured.
What Jev is
Jev (typesafe/jev-1.13) is TypeSafe’s structured decision model. It doesn’t chat. You give it a state, a question and a set of options, and it returns a choice with a probability for each option. On OpenRouter it sits behind a separate decisions endpoint instead of chat completions.
POST /api/alpha/decisions
{
"model": "typesafe/jev-1.13",
"state": { "path": "saleor/app/utils.py", "content": "…" },
"questions": { "relevant": {
"type": "choice",
"instructions": "Question: \"How are taxes calculated?\". Is this file part of this feature, i.e. a file a developer would likely need to edit when this feature changes? Files that only mention it in passing do not count.",
"criteria": {
"yes": "The file is part of this feature.",
"no": "The file is not part of this feature, or only mentions it in passing."
}
}}
}
→ { "choice": "yes", "probabilities": { "yes": 0.66, "no": 0.34 } }The probability is what makes it useful. It turns Jev into a dial: keep every file with a “yes” probability of at least 0.5 for more recall, or at least 0.7 for more precision. Input costs $0.042 per million tokens, output is free, and the context window is 32k tokens, so Jev judges one file per call.
How we tested
We started from the six Saleor questions in the grep study. To get past 5,000 candidate files we added five Go projects from the Kubernetes ecosystem, with five questions each:
- kubernetes: the scheduler, Dynamic Resource Allocation, API Priority and Fairness, the Job controller, admission
- istio: xDS push, workload certificates and mTLS, ambient mode, sidecar injection, AuthorizationPolicy
- argo-cd: sync, cluster management, ApplicationSets, RBAC, the repo server
- cert-manager: ACME, the Ingress and Gateway shim, the webhook, the Vault and Venafi issuers, renewal
- etcd: watch, Raft apply, leases, snapshots, membership
Every question goes through the same pipeline as the grep study:
- Candidates. Grep the question’s terms (for example
vault|venafi) over non-test, non-generated source files. The judges only ever see grep hits. - Answer key. A file is a right answer if it was changed in at least three commits since January 2023 whose subject matches the topic (for example
renew|expir|reissu). The key comes from the commit log, never from a model. - Hubs. Files that land in the keys of three or more questions in the same repo, like shared
types.gofiles, are left out of the scoring. - Three depths. Each judge sees one of: D1, the matching grep lines; D2, an outline of functions and types; D3, the full file.
| Count | |
|---|---|
| Codebases | 6 (Saleor plus five Go projects) |
| Questions | 31 |
| Grep candidates per depth | 5,075 (4,062 new, 1,013 Saleor) |
| Right files that grep found | 87%, the ceiling for any judge |
The two judges. Jev judged every candidate at every depth, one file per call: 15,225 decisions. Opus 5.5 read the same packets in batches of up to about 400,000 characters, with the benchmark’s original instruction, and returned a list of relevant paths. Opus ran on a subset of questions because of cost and time. Every comparison below uses only the question and depth pairs that both judges completed, on exactly the same files.
Results
Recall is the share of right files a judge found. Precision is the share of its picks that were right. These totals add up file counts across all questions at each depth.
| Depth | Grep gave | Right answers | Right answers in grep | Opus picked → right | Opus recall / precision | Jev ≥0.5 picked → right | Jev ≥0.5 recall / precision | Jev ≥0.7 picked → right | Jev ≥0.7 recall / precision |
|---|---|---|---|---|---|---|---|---|---|
| D1 grep lines (20 sets) | 2,170 | 392 | 339 | 470 → 183 | 47% / 39% | 791 → 228 | 58% / 29% | 643 → 208 | 53% / 32% |
| D2 outline (18 sets) | 2,007 | 422 | 367 | 384 → 165 | 39% / 43% | 518 → 210 | 50% / 41% | 420 → 187 | 44% / 45% |
| D3 full file (16 sets) | 1,577 | 338 | 289 | 397 → 151 | 45% / 38% | 491 → 174 | 51% / 35% | 408 → 154 | 46% / 38% |
Both judges saw the same files within each row. The question sets differ between rows, so compare judges within a depth, not across depths.
What the numbers say:
- Jev finds more of the right files at every depth. With the 0.5 threshold it found 228 right files at D1 against Opus’s 183, and 210 against 165 at D2.
- Opus is more selective. Its picks are right more often on grep lines (39% against 29%). Jev pays for its extra recall with extra picks.
- With the 0.7 threshold, the two meet. On outlines, Jev is slightly ahead on both measures: 44% recall and 45% precision against Opus’s 39% and 43%. On full files they are level: 154 right files out of 408 picks for Jev, 151 out of 397 for Opus.
- Full files don’t buy much. Opus’s recall on full files (45%) is about the same as on grep lines (47%), for about four times the input. For both judges, the grep lines or the outline are enough to decide.
Cost
On the same files, the difference is two orders of magnitude.
| Judge | Scope | Input tokens | Cost | Per file | How we measured it |
|---|---|---|---|---|---|
| Jev | D1 + D2, matched set (4,177 files) | 3.96M | $0.17 | $0.00004 | OpenRouter billing, per call |
| Opus 5.5 | D1 + D2, matched set (4,177 files) | ≈5.2M | ≈$16–21 | ≈$0.004–0.005 | Token counts × list price; a paid OpenRouter run came to $0.0039 per file |
| Jev | All 31 questions, D1 to D3 (15,225 decisions) | 33.0M | $1.39 | $0.00009 | OpenRouter billing |
Jev’s median latency was 0.39 seconds per file with 8 requests in flight. Opus took 20 to 120 seconds per batch.
Jev judged every file in the benchmark, at all three depths, for 16 to $21.
A closer look: Saleor tax
Grep returned 107 files for “How are taxes calculated?”. 36 files are on the answer key, and 19 of those are among the grep hits. On grep lines, Jev with the 0.5 threshold picked 32 files and got 11 right. Opus picked 23 and got 9.
The disagreements are instructive. Jev reliably kept the GraphQL tax layer (graphql/tax/types.py, graphql/core/types/taxes.py, tax_configuration_update.py), which the batch judges tended to skip. It reliably rejected checkout/base_calculations.py and order/base_calculations.py. Those files compute prices before tax, and they’re on the key because tax commits keep touching them. A human could argue either way. The answer key says they count.
Recall per question
| Repo | Question | D1 Opus | D1 Jev ≥0.5 | D1 Jev ≥0.7 | D2 Opus | D2 Jev ≥0.5 | D2 Jev ≥0.7 | D3 Opus | D3 Jev ≥0.5 | D3 Jev ≥0.7 |
|---|---|---|---|---|---|---|---|---|---|---|
| argo-cd | argoappset | 40% | 67% | 67% | 40% | 67% | 65% | 44% | 67% | 67% |
| argo-cd | argocluster | 71% | 76% | 71% | 53% | 71% | 71% | 59% | 76% | 76% |
| argo-cd | argorbac | 64% | 73% | 73% | 64% | 73% | 64% | 64% | 73% | 73% |
| argo-cd | argorepo | 33% | 27% | 27% | 33% | 33% | 27% | – | – | – |
| argo-cd | argosync | 38% | 75% | 62% | 38% | 50% | 50% | – | – | – |
| cert-manager | cmacme | 95% | 100% | 100% | 85% | 95% | 80% | 95% | 100% | 80% |
| cert-manager | cmrenew | 83% | 83% | 83% | 83% | 83% | 83% | 83% | 83% | 67% |
| cert-manager | cmshim | 62% | 85% | 69% | 46% | 38% | 31% | 62% | 54% | 31% |
| cert-manager | cmvault | 92% | 100% | 92% | 92% | 92% | 92% | 92% | 92% | 75% |
| cert-manager | cmwebhook | 30% | 35% | 30% | 26% | 26% | 22% | 30% | 26% | 22% |
| etcd | etcdlease | 100% | 100% | 100% | 75% | 75% | 75% | 75% | 100% | 75% |
| etcd | etcdmember | 20% | 20% | 20% | – | – | – | 20% | 0% | 0% |
| etcd | etcdraft | 62% | 75% | 62% | 62% | 62% | 62% | 62% | 62% | 62% |
| etcd | etcdsnapshot | 64% | 82% | 73% | 55% | 73% | 73% | 64% | 73% | 64% |
| istio | istioambient | – | – | – | 30% | 44% | 37% | – | – | – |
| istio | istioauthz | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% |
| istio | istioinject | 20% | 30% | 30% | 30% | 20% | 20% | – | – | – |
| istio | istiomtls | 57% | 67% | 57% | 67% | 57% | 52% | – | – | – |
| saleor | discount | 32% | 45% | 36% | 18% | 32% | 25% | 23% | 32% | 27% |
| saleor | giftcard | 90% | 90% | 90% | – | – | – | 90% | 90% | 90% |
| saleor | tax | 25% | 31% | 31% | – | – | – | 28% | 25% | 19% |
| Average | 58% | 67% | 63% | 54% | 60% | 56% | 61% | 65% | 57% |
Averages weight every question equally. A dash means that pair wasn’t part of the matched comparison.
Precision per question
| Repo | Question | D1 Opus | D1 Jev ≥0.5 | D1 Jev ≥0.7 | D2 Opus | D2 Jev ≥0.5 | D2 Jev ≥0.7 | D3 Opus | D3 Jev ≥0.5 | D3 Jev ≥0.7 |
|---|---|---|---|---|---|---|---|---|---|---|
| argo-cd | argoappset | 81% | 60% | 67% | 81% | 71% | 76% | 83% | 64% | 67% |
| argo-cd | argocluster | 67% | 37% | 39% | 69% | 52% | 60% | 62% | 57% | 62% |
| argo-cd | argorbac | 54% | 38% | 42% | 58% | 47% | 54% | 54% | 36% | 42% |
| argo-cd | argorepo | 38% | 33% | 50% | 33% | 42% | 67% | – | – | – |
| argo-cd | argosync | 21% | 18% | 18% | 14% | 12% | 17% | – | – | – |
| cert-manager | cmacme | 28% | 21% | 23% | 30% | 25% | 24% | 28% | 24% | 24% |
| cert-manager | cmrenew | 36% | 20% | 28% | 42% | 31% | 36% | 31% | 22% | 25% |
| cert-manager | cmshim | 73% | 58% | 69% | 67% | 62% | 57% | 73% | 64% | 57% |
| cert-manager | cmvault | 58% | 50% | 58% | 65% | 65% | 69% | 55% | 52% | 56% |
| cert-manager | cmwebhook | 39% | 29% | 32% | 33% | 29% | 33% | 33% | 32% | 29% |
| etcd | etcdlease | 13% | 6% | 8% | 17% | 10% | 12% | 12% | 12% | 11% |
| etcd | etcdmember | 5% | 4% | 5% | – | – | – | 5% | 0% | 0% |
| etcd | etcdraft | 42% | 30% | 33% | 42% | 42% | 50% | 38% | 38% | 62% |
| etcd | etcdsnapshot | 37% | 21% | 27% | 43% | 30% | 40% | 32% | 29% | 33% |
| istio | istioambient | – | – | – | 50% | 54% | 54% | – | – | – |
| istio | istioauthz | 22% | 13% | 14% | 27% | 24% | 29% | 20% | 17% | 18% |
| istio | istioinject | 17% | 13% | 20% | 19% | 12% | 18% | – | – | – |
| istio | istiomtls | 32% | 22% | 27% | 34% | 32% | 38% | – | – | – |
| saleor | discount | 82% | 70% | 69% | 77% | 77% | 78% | 81% | 72% | 78% |
| saleor | giftcard | 20% | 12% | 14% | – | – | – | 18% | 18% | 20% |
| saleor | tax | 39% | 34% | 39% | – | – | – | 36% | 33% | 30% |
| Average | 40% | 29% | 34% | 44% | 40% | 45% | 41% | 36% | 38% |
Precision is understated for every judge, because the commit-log key misses some files that are clearly relevant.
What grep gave each judge (D1)
| Repo | Question | Grep gave | Right answers | Right answers in grep | Opus picked | Opus right | Jev ≥0.5 picked | Jev ≥0.5 right |
|---|---|---|---|---|---|---|---|---|
| argo-cd | argoappset | 92 | 43 | 35 | 21 | 17 | 48 | 29 |
| argo-cd | argocluster | 106 | 17 | 17 | 18 | 12 | 35 | 13 |
| argo-cd | argorbac | 60 | 11 | 8 | 13 | 7 | 21 | 8 |
| argo-cd | argorepo | 106 | 15 | 11 | 13 | 5 | 12 | 4 |
| argo-cd | argosync | 131 | 8 | 8 | 14 | 3 | 33 | 6 |
| cert-manager | cmacme | 134 | 20 | 20 | 67 | 19 | 94 | 20 |
| cert-manager | cmrenew | 61 | 6 | 5 | 14 | 5 | 25 | 5 |
| cert-manager | cmshim | 34 | 13 | 13 | 11 | 8 | 19 | 11 |
| cert-manager | cmvault | 35 | 12 | 12 | 19 | 11 | 24 | 12 |
| cert-manager | cmwebhook | 124 | 23 | 22 | 18 | 7 | 28 | 8 |
| etcd | etcdlease | 127 | 4 | 4 | 31 | 4 | 63 | 4 |
| etcd | etcdmember | 108 | 5 | 5 | 21 | 1 | 25 | 1 |
| etcd | etcdraft | 130 | 8 | 8 | 12 | 5 | 20 | 6 |
| etcd | etcdsnapshot | 82 | 11 | 11 | 19 | 7 | 43 | 9 |
| istio | istioauthz | 96 | 5 | 4 | 18 | 4 | 30 | 4 |
| istio | istioinject | 188 | 10 | 10 | 12 | 2 | 23 | 3 |
| istio | istiomtls | 168 | 21 | 21 | 38 | 12 | 65 | 14 |
| saleor | discount | 173 | 114 | 97 | 44 | 36 | 73 | 51 |
| saleor | giftcard | 108 | 10 | 9 | 44 | 9 | 78 | 9 |
| saleor | tax | 107 | 36 | 19 | 23 | 9 | 32 | 11 |
| Total | 2,170 | 392 | 339 | 470 | 183 | 791 | 228 |
Limits of this study
- The answer key is noisy. “Changed in three topic commits” catches files that move with a feature, not a human’s judgment of relevance. Both judges are scored against the same key, so the comparison is fair, but absolute precision is understated.
- Opus covered a subset. For cost reasons Opus judged 20 question sets at D1, 18 at D2 and 16 at D3, and none of the five Kubernetes questions. Jev’s results on the rest are in the dataset but aren’t part of the head-to-head.
- Batches against single files. Opus saw up to a few hundred files at once and could compare them. Jev saw one file at a time. Jev’s input tokens also include its instruction, repeated on every call.
- The threshold was picked after scoring. We report 0.5 and 0.7 side by side. Choosing the better one after the fact flatters Jev slightly.
- Minor run issues. Two Opus batches on one Argo CD question skimmed part of their packets instead of reading them in full. 70 of 5,075 full files were cut to fit Jev’s 32k-token window, 13 of them in the matched D3 set. No Jev call failed.
What this means
- Keep grep for recall. It already surfaces 87% of the right files. The judging step decides quality, and it’s where the money goes.
- A small decision model can replace a frontier model as the judge. On the same grep candidates, Jev finds more of the right files, and with the 0.7 threshold it matches Opus 5.5’s precision, for about 1% of the cost.
- Show the judge less. Grep lines or an outline gave the same decisions as full files, for both judges. That cuts Jev’s cost further and removes almost all truncation.
- Use the probability. One run gives you a recall and precision dial. Lower the threshold when missing a file is costly. Raise it when the context window is tight.
In practice, a coding agent can grep, send each hit’s matching lines to Jev, keep everything above 0.7, and hand the survivors to the frontier model. That model then reads the files that matter instead of spending its budget deciding which ones do.
The ceiling doesn’t move, though. A judge can only keep what grep found, and grep still can’t follow inheritance or code that never says the name. That’s the gap we described in the grep study, and a cheaper judge doesn’t close it.
Reproducing it. Repos are pinned at Saleor 58e2eb7, kubernetes 8f58e1b, istio a464d0e, argo-cd d89c605, cert-manager fd7d37b and etcd 7583cc6. The candidates, answer keys and packets are rebuilt from those commits, and every judge’s picks are scored with the same script.
ByteBell checks AI-written code against your whole codebase, inside your network. It indexes every repository once, knows what calls what and what inherits from what, and hands the model a short, connected set of files with the line behind every claim. Learn more at bytebell.ai.