I Didn't Become a Developer to Review AI Slop
Updated on October 2, 2026.
But lately, that's exactly what the job feels like.
My PR queue fills with work that, yes, technically compiles. The summary sounds plausible. It might even have some tests. Then I open the diff, and the real work starts.
What was the change supposed to do? Did anyone actually run the flow? Why is this helper duplicated six times? Is this actually fixing a bug, or did the AI just run around in circles and call it done?
AI made it effortless for anyone on my team (and yours) to create code, but it didn't make that code trustworthy.
Stack Overflow's 2025 Developer Survey found the most common frustration with AI tools is output that's "almost right, but not quite." Sonar's 2026 State of Code report found that 96% of developers don't fully trust AI-generated code, and 38% say reviewing it takes more effort than reviewing human-written code.
I believe it. AI code looks fine, so you have to really dig in to figure out what it's actually doing. Straight-up bad code is way easier to reject.
I'm annoyed. Maybe you are, too.
A PR is... almost too cheap now
AI agents can spin up branches from Jira tickets, patches from Slack threads, or even full PRs from a bug report before anyone agrees the bug is real. It's honestly a pretty awesome world. Our open-source repo, Agent-Native, merged more than 600 PRs in the last week of September alone.
But developers aren't the only ones using these tools. PMs will prototype the feature they've been trying to explain for three sprints, mostly with vague, unhelpful hand waves. Designers will tweak UX flows and fix layouts that keep getting deprioritized. Marketers will update landing pages and forms. (Constantly.) Support will patch the customer pain points they know best.
And all of that is a win. Small fixes shouldn't sit in backlog hell waiting for an engineer who happens to know that part of the code. Product knowledge should be turning into working software faster.
But the easier it gets to open a PR, the more PRs developers have to review. And a PR nobody can trust isn't worth much.
And trust is still really expensive
AI is really good at writing code. At the end of September, I had Opus 5.5 build a way for outside agents like Claude Code and Cursor to read data from our Analytics app directly. The PR touched 58 files, came with tests, and the tests passed. Honestly, it was good work.
Then I had GPT-6.1 Sol review the diff, and it found four real problems. My favorite: an agent could call a free tool that only lists what's available, and the app would count that as proof it had read customer data. Another was a test that checked which tools we meant to expose instead of which ones the server actually served. (It served more.) And a fix for one bug meant that if two teammates opened the same old dashboard at the same moment, the second one got "not found."
None of that showed up in the tests. We fixed all of it before merging, but only because someone went looking.
Writing code and writing trustworthy, scalable code are two different things. A model can write a diff, explain it, and run some happy-path tests. But someone's still accountable for the stuff that actually matters:
- Did this code actually fix the stated problem?
- Did the author really understand the system, or is this tech debt for later?
- Is the diff bigger than it needs to be? (Almost definitely.)
- Does this fix break some other flow that a single user would notice in five minutes?
- Does the UI actually work for real users in real browsers?
- Is this a fix for the root problem, or just a bandaid?
- Is this security tradeoff acceptable?
Those are trust questions, and right now they all land on you and me, the developers. @richiemcilroy put it well in a viral video:
The research says the same thing. LinearB's 2026 benchmarks found AI PRs wait 4.6x longer for review and get rejected way more often than human-written ones. And in March, METR had four maintainers from scikit-learn, Sphinx, and pytest re-review 296 AI-written PRs that had already passed SWE-bench's automated tests. They wouldn't have merged about half of them.
Those PRs came from 2024 and 2025 agents, so you could argue newer models have fixed this. My Analytics PR passed its tests too.
The models really do keep getting better. But typing code into files was never the hard part of software. The hard part is knowing what should change, what shouldn't, and when a patch that technically fixes the problem is going to haunt your team for the next six months. That's what your taste and context are for, and it's where your attention should go. Not on rediscovering the basics after a PR is already in your queue.
Developing feels bad right now
Even though AI tools are making everyone more productive, being the bottleneck feels terrible. Everyone else gets to do more than they ever could, because suddenly code is open to them.
Developers mostly experience the hype as incoming review debt. Your day turns into reverse-engineering what an agent or a teammate was trying to do, then betting your afternoon on whether the diff is safe to keep.
The AI gets to do the fun part. You get to be a robot.
You're not useless, obviously. Your judgment matters more now than it did before any of this.
But the workflow spends that judgment terribly. The scarcest thing on most teams is an experienced engineer's attention, and we're pointing it at mystery diffs, bloated patches, and code that only looks correct.
So yeah, it's boring. Yeah, it's frustrating. When someone says "now everyone can ship code," what you and I hear is "now everyone can create work for us."
No wonder we're all burnt out.
Locking down the repo solves the wrong problem
So, what do we do? The obvious reaction is to lock up the repo. Devs only.
And I get that. You're the one who gets paged at 2am when prod goes down. Being protective of the code isn't elitism. You just have a memory.
GitHub even shipped a version of it. Since June, repo admins can cap how many open PRs someone without write access can have. If you maintain an open-source project that's getting buried in drive-by AI PRs, fair enough. But a cap only limits how many PRs you get. The ones under the cap aren't any easier to trust, and the PM on your own team already has write access anyway.
The other reflex goes the opposite way: let a bot approve the PRs. That shipped too. Since September, Copilot's code review can approve a PR, and its approval counts toward your required reviews. (It's in preview and off by default.) Required reviews exist so a second person looks at the change. A bot approval satisfies the rule without one.
And honestly, cross-functional PRs are a lot of what we've wanted for years: product knowledge turning into small fixes without waiting on an engineer's calendar.
What hasn't changed is how we take PRs in. Teams still treat a PR like a dev-to-dev handoff: here's the diff, here's the description, good luck. That worked when the author was another engineer with the same local context, the same testing habits, and the same gut sense of what reviewers need.
Now the author might be a designer who's never opened the test suite, and that's fine. Designers, PMs, marketers, and support usually have the best user context on the team, because they're closest to the problem. They just don't know what you need to judge risk. And when AI wrote the implementation, even the person opening the PR might not know everything it changed.
Raise the bar for evidence on PRs
No developer should have to reverse-engineer a generated PR from scratch. Every PR should show up with:
- What it's supposed to do, in plain words.
- A diff no bigger than it needs to be.
- Tests, and whether they passed.
- Proof it works in a real browser: screenshots, a replay, or a recording.
- Console and network logs if anything failed.
- What it didn't test, and what's still a judgment call.
But that's the problem. We say we want PMs, designers, marketers, and support to contribute directly, and then we expect them to act like senior engineers before we'll look at their work. A PM shouldn't have to scope a tight diff, and a designer shouldn't have to read network traces.
The entry bar needs to stay low. The review bar needs to go up.
You can have both if AI does the prep work on every PR before an engineer sees it. The contributor brings the product context: what hurts, why it matters, what good looks like. They work with an agent to open the PR. Then automation checks the work, collects that evidence, and sends problems back to the same branch, so the contributor can fix them without turning into a release engineer.
And no developer has to be the first person to find out the button doesn't work.
Review bots finally run the code
When I first wrote this in May, most AI review tools just read the diff and had opinions about it. That changed fast. Copilot's code review can now run your builds and tests, and it has Playwright switched on by default. Cursor's agents attach screenshots and video to their PRs. OpenAI announced automatic Codex review at DevDay.
Good. Running the code should be the minimum. What I want to know is whether the reviewer hands back the evidence from that list above: what it clicked, what broke, and what it skipped.
Here's how that works in our own repo. Every PR in Agent-Native gets an automatic visual recap from Agent-Native Plans: an image of what changed, linked to an interactive version, so you know what you're looking at before you open the diff. Of our last 50 merged PRs, 49 had one, and the 50th came from our release bot. (I wrote more about recaps in Stop watching your agent work.) Our review bot reviewed all 50. And for riskier changes, like that Analytics PR, a second model reviews the first one's work before anything merges.
None of this does the review for us. On that same Analytics PR, our bot caught a paid request that could get sent to the wrong server, and it also flagged two things that turned out to be fine. I still made the calls on what to fix and how. But I made them with the evidence in front of me, instead of rebuilding it from the diff myself.
If you want recaps on your own PRs, Plans is free and open source, and the workflow that posts ours is in the repo.
Don't become the human merge queue
AI made it easy for anyone to open a PR. Locking people out won't fix that, and neither will letting a bot hit approve. What helps is getting every PR to show up with enough proof that you can spend your brain on judgment calls instead of playing detective.
We put ours together one piece at a time while maintaining Agent-Native, and I wrote up the whole setup in Build an agentic software factory, starting with one bug. If you want to see it running, try Agent-Native or browse the source.
Whatever you use, deal with this before you turn into your company's human merge queue.