All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly. I also still manually review code, and usually also catch issues, but the severity of what I find shrinks ever further as agents get better.
The sprawling code comments are becoming the most draining part of code review though, that's really killing me from the inside.
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly.
No, none of today's AI would give you enough signal around "should this thing be built in the first place" nor if it's the correct solution on a high level.
They don't understand why you are doing what you are doing, and even if you explain it, they still don't actually understand the motivation and lots of other things.
You'll get them to do guesses and pretend they actually know how to prioritize and will tell you it makes lots of sense, whatever they come up with. But try following it blindly and you'll see where you end up.
This is why "one agent + one good developer" beats "thousands of agents working in a swarm" still today.
I don’t think I claimed agent reviews to be a panacea. It’s a tool that can help you lower the review pressure in companies working with agentic coding tools.
> What really important things are human reviews catching in your org?
Another person said:
> 1. Whether the thing should be done in the first place
And you replied:
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly
Which as I noted, is very far from the truth. I neither claimed that you said "agent reviews are a panacea", but when you claim "AI can solve all those things" and two of the first items cannot be addressed by AI (today), then I'm rebuking those specific things, not some other general point you implicitly made.
It is mostly useless at figuring out if "should this thing be built in the first place" and "if it's the correct solution", and mostly cannot help at all with those things.
Where "mostly" means kind of what it says but also not really.
The AI review are still quite far from having the same level of critical thinking and high level knowledge of your application, what you have done in the past and want to do next etc.
If you don't master this for your own project, what's even the point of your job.
Tbh I just go through openrouter because it's easy, so you can just go down the list in that and see who sounds good. I honestly don't spend much time at it.
The Turing Test is about being able to figure out if “someone” is a computer during a short conversation, not “millions of articles”.
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
We don't train current LLMs to mimic the average human's writing style. We train it to be smart, helpful, and knowledgable, too knowledgeable for a human. We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.
I just feel more and more like the effort invested in manual reviews is not worth it
reply