← All posts

We Measured Grep on Real Codebases. Here's What Actually Breaks.

We measured how coding agents search with grep on Home Assistant and Saleor: about 150 model runs and 13 million tokens. Grep finds the files, then buries them in noise, and it can't follow inheritance or code that never says the name. A bigger model doesn't fix it.

We Measured Grep on Real Codebases. Here's What Actually Breaks.

We Measured Grep on Real Codebases. Here’s What Actually Breaks.

Coding agents find code the way we did in 2005: grep, find, read the file, repeat. Claude Code does it. Cursor does it. Devin does it. The argument for it is simple. There’s no index to go stale, and a smart model can recover from a bad search.

We wanted to know if that holds up as AI writes more of the code and pull requests grow to 10 or 30 files. So we stopped arguing and measured it. Two real open-source codebases, about 150 model runs and roughly 13 million tokens.

The short answer: grep is a real limit, but not the one people usually point to. It rarely fails to find files. It buries the right ones in noise, and it can’t follow code that is connected without sharing a name. A smarter model doesn’t fix either problem.

How we tested

We used two Python codebases with long histories:

  • Home Assistant, 28,213 files. Big, repetitive, thousands of integrations built on the same base classes.
  • Saleor, an e-commerce backend with 4,675 files and 22,663 commits.

The hard part of a test like this is the answer key. If you use grep to decide which files are “right”, grep will always look perfect. So we took answer keys from git history instead: the files people actually changed in commits about a topic. For example, files touched by at least three commits about payments, refunds or transactions since 2023. Files that almost every commit touches, like shared schema modules, we left out of the scoring.

Then we split the work into the two things an agent really does:

  1. Find candidates. Grep for the topic words, or ask an embedding model for the closest files.
  2. Decide. A model reads the candidates and keeps the relevant ones. We gave it three reading depths: the grep lines only, a file outline, or the whole file.

Every judge got the same instructions and read only its packet. We ran Sonnet for most judges and Opus where the model itself was the question. We also looked at 240 real commits that touched 10 to 30 source files, to see what writing and reviewing a normal PR looks like through grep.

Grep finds the files, then buries them

We asked “How does payment work?” in Saleor and ran the ten searches an agent would plausibly run. Grep found 43 of the 44 files on the answer key. It also returned 376 files and about 332,000 tokens to get there. 89% of what it handed back was wrong.

The first search alone, a plain grep for “payment”, produced about 94,000 tokens. That’s most of a typical context window, spent before the agent has read a single file properly.

The same pattern shows up in ordinary code changes. We took real commits, started from the file with the biggest edit, and grepped every distinctive name touched there. Grep found 83% of the other files the commit changed in Saleor and 94% in Home Assistant. But it returned a median of 60 files in Saleor and 613 in Home Assistant to find the two to four that mattered.

So the usual complaint, “grep can’t find it”, is mostly wrong. Grep finds it. The problem is everything else it finds.

Where grep really breaks: code that never says the name

Grep matches text. Code is connected by more than text.

Inheritance. In Home Assistant, AcaiaSensor inherits from AcaiaEntity, which inherits from CoordinatorEntity. The sensor file never contains the word “CoordinatorEntity”. We built the real inheritance tree with Python’s parser and found this is normal, not rare. 62% of the files built on CoordinatorEntity never mention it. Ask “what breaks if I change CoordinatorEntity?” and grep sees 761 files. The real answer is about 1,734. Closing that gap by hand takes around 650 follow-up searches, up to six levels deep.

One name, many implementations. A light turns on through a single line, light.async_turn_on(...). 635 files define an async_turn_on. Grep can list all of them. It can’t tell you which one runs.

Concepts that don’t use the word. For “How are taxes calculated?”, grep found only 53% of the tax files. The missing ones were discount models and checkout mutations. They change every time tax logic changes, but they never say “tax”.

Real 10 to 30 file PRs: a third to half of lookups aren’t clean

New code leans on old code. A typical PR in our sample touched 14 to 17 files and used 20 to 29 existing functions and classes. For each of those, we checked what happens when you grep its name.

SaleorHome Assistant
Name defined in one place (grep lands on it)68%51%
Defined in 2 to 5 places (model must work out which)28%25%
Defined in more than 5 places (grep can’t tell)4%23%
PR files reachable by grepping the names it changes64%46%

In other words, 32% to 49% of the lookups in a real change are ambiguous or unresolvable by name. Home Assistant is worse because every integration defines the same method names.

How you search matters enormously. Grepping every name a PR uses returned a median of 1.65 million tokens in Saleor and 3.65 million in Home Assistant. A precise definition-only search for the same names cost about 2,000 and 9,500 tokens.

Review is the weaker side. If a reviewer follows the names a PR adds or changes, grep reaches only 64% of the PR’s own files in Saleor and 46% in Home Assistant. That’s before asking what the change breaks elsewhere.

The model is the second filter, and it filters by file name

This was the finding we didn’t expect. Across six questions (payment, discounts, shipping, tax, gift cards, stock), grep put 83% of the right files in front of the model. The model kept 43%. It threw away about half the good files it was shown.

Why? Look at what the model actually receives from grep: a path, a line number and one matching line. It hasn’t opened the file. It decides mostly from the path.

We checked this on a subagent’s picks for the payment question:

Has a payment-related word in the path
Right files it found89%
Right files it missed25% (2 of 8)
Wrong files it picked85%

It found the right files with good names. It missed right files with neutral names, like site/models.py. And it picked dozens of wrong files because their names sounded right. A file with a neutral name that holds key logic fails both filters. That’s exactly the file a code reviewer can’t afford to miss.

Reading more doesn’t help. Less noise does. A bigger model doesn’t.

The obvious fix is to let the model read the whole file before deciding. We tried it on all six questions.

What the model readRight files keptPicks that were rightTokens spent per question
Grep lines only43%48%309,000
File outline (names only)39%50%156,000
Whole files45%44%1,049,000

Whole files cost about 3.4 times as much and changed almost nothing. The cheapest option, just function and class names, was the most efficient per right file found. The model mostly decides from names and structure, not from the code body.

Next we tested the context itself. We gave Opus the exact same grep results, either all at once (35,000 to 93,000 tokens) or split into batches of about 8,000.

SetupRight files kept
Sonnet, one big batch51%
Opus, one big batch45%
Opus, small batches55%

A smaller context raised recall by about 10 points. For the biggest batch, stock, it went from 26% to 43%. So a context stuffed by grep really does cost accuracy.

A bigger model didn’t help at all. Opus with one big batch did no better than Sonnet. The limit here is what the model is given, not how smart it is.

Subagents and embeddings

Subagents move the cost, they don’t remove the problem. The popular fix is to run the searching in a separate agent and pass back only the results. We sent one at “How does payment work?”. It found 82% of the payment files, missed all four files that never say “payment”, and handed back 111 files. Reading them would cost about 474,000 tokens. The main agent’s context stays clean, but whatever grep couldn’t see stays unseen.

A small embedding model isn’t the answer either. We ran all-MiniLM-L6-v2, a small general-purpose text model, over every file. Its results were about nine times cheaper to read than grep’s, and its picks were slightly more accurate. But only 46% of the right files made its top 60, against 83% for grep. A model trained on code would likely do better, and we didn’t get to test one. Even then, embeddings only improve the first step. The second step, the model’s judgment, was the bigger loss in every setup we ran.

What we got wrong going in

We build a code verification tool, so we started with opinions. Several didn’t survive.

  • “Grep can’t find a class used in a hundred places.” It can. A plain grep for DataUpdateCoordinator in Home Assistant returns 152,000 tokens from 1,752 files. A grep for where it’s defined returns about 100 tokens and the right file.
  • “Better models will fix it.” Opus did no better than Sonnet when both got the same crowded context.
  • “The model needs to read more to decide.” Whole files cost 3.4 times as much for the same accuracy.
  • “AI means ten times more code.” The best field data we found shows about two times more output per developer, not ten.
  • “Code graph tools are just fancy grep.” Some are, if they only parse text. Tools built on compiler-resolved references are a different thing, and they’re the ones that follow inheritance.

What this means for code review

Put the pieces together and the picture is fairly clear.

  1. Grep-based agents find most of the code, but they pay for it in noise. Getting the two to four files a change needs means wading through 60 to 600.
  2. A third to half of what a real PR touches is connected by inheritance, dynamic calls, or files that simply change together. Grep can’t follow those links, and a subagent can’t either.
  3. The model then filters what grep found mostly by file name, and throws away about half the right files.
  4. Giving it more to read doesn’t fix that. Giving it less noise helps a little. Giving it a bigger model doesn’t help.

The bottleneck isn’t search and it isn’t intelligence. It’s that neither grep nor the model knows how the code is connected. A review tool that already knows what calls what, what inherits from what, and what has changed together in the past can hand the model a short, connected set of files instead of a pile. That’s the gap every setup in this test failed at.

Even perfect retrieval doesn’t make raw files cheap. The 44 right files for “How does payment work?” cost 337,000 tokens to read. For questions like that, the right unit is a prepared description of each piece and how the pieces link, not the file itself.

Limits of this study

This is a first pass, not a paper. Keep these in mind before quoting a number.

  • Python only, two repos. Typed languages with better tooling may look different.
  • The answer keys are noisy. They come from commit history, so some files the judges picked as “wrong” were clearly right. checkout/webhooks/calculate_taxes.py is part of tax calculation but wasn’t changed often enough to make the key. Precision is understated for every method. Comparisons between methods are fair because all of them were scored against the same key.
  • Mostly single runs. The context-size test covered three questions. Treat 45% to 55% as a direction, not a precise figure.
  • Token counts are estimates, characters divided by four, except where a model reported its own usage.
  • Not tested: code-trained embedding models, compiler-grade indexes, and fast classifier models as the judge. Those are next.

ByteBell checks AI-written code against your whole codebase, inside your network. It indexes every repository once, knows what calls what and what inherits from what, and hands the model a short, connected set of files with the line behind every claim. Learn more at bytebell.ai.

All posts