Yesterday, I had again one of those experiences that are funny not funny. I am currently building an app with Claude Code and working my way into specification-driven engineering. I have a team of several virtual engineers and advisors. So yesterday, I run an architectural review, led by the lead engineer. Afterwards, we ran a blind test with a new reviewer role – including role description, and review description and goals. Both reviews produced overlapping findings in some places and complementary findings in others.
- Me to the reviewer: How is it possible that you missed things the lead engineer found?
- Reviewer: I have to admit that I have not read the entire repository. I have skipped parts of it. Sorry.
- Me: How can that happen??
Then I brought this finding—the meta-analysis of the review itself—back to the lead engineer.
- Me to the Lead Engineer: How do we classify this? How can we improve the review going forward?
- Lead Engineer: First of all, it was good and honest feedback from the reviewer. – But I also have to admit that I have not read the entire repository.
??? … Come on guys.
Of course, honesty is better than pretending a review was complete. But if someone reviews a repository without reading the full repository – or at least without following a clearly defined coverage strategy – how can I trust? My current answer is: More discipline, better „job role description“. Although Claude produces acceptable code, it doesn’t necessarily mean that engineering is applied the way we are used to in a human, classical engineering team.
And that is where there is still a lot to do in AI engineering: not only in what AI can find, but in how we structure the work around it – in this case the review. What was reviewed? What was skipped? What does the result actually cover? And when do we know that a review is complete enough to rely on it?
There still is some work to do. … All this must be figured out before giving it to fully autonomous agents.



