Ask an agent to review the code it just wrote and it will find something. The findings will be well phrased and ranked by severity, and they will almost never include the one that matters: that it built the wrong thing.

An agent that has spent an hour on a task has stopped reading the requirement as a requirement. It reads it as a description of what it built, because every ambiguity has already been resolved in favour of the code that exists. So my answer to the title is no. The more useful question is what counts as independent, and my first answer to that was wrong.

Fresh context is not enough

The obvious fix is a clean session: new conversation, paste in the diff, ask for a review. I ran that for a while. It catches sloppiness, which is worth having.

What it misses is the confident wrong interpretation. The fresh session is still the same model, with the same training and the same instincts about how to read an unclear requirement. Hand it the same ambiguity and it will resolve it the same way. Clearing the memory of a decision does nothing about the tendency that made it.

How I run it now

Claude Code orchestrates and keeps the full context: the refinement conversation, the constraints, the decisions taken and the options rejected. That context took a long time to build and I want it intact for the rest of the work.

At review time it starts a separate CLI process running a model from a different vendor, and gives it two things: the diff, and the acceptance criteria exactly as they were agreed before implementation began. The reviewer gets no summary of the implementation and no explanation of the choices behind it.

That omission is deliberate, and it is the step most people would reverse. More context feels like it should produce a better review. Here it does the opposite, because the orchestrator's explanation is a persuasive account written by the party that wants the work accepted.

The reviewer returns met, not met or cannot tell against each criterion, with evidence. Cannot tell turns out to be the most useful of the three, and it is a verdict I have never seen a model give about its own work.

What the evidence does and does not show

I have run the reviewer position with Codex, with Opus, and with a model reserved for security-sensitive changes. On one release, routing had leaned away from Codex, nine reviews to seven. Codex still found the most serious defect in the release. I made it the default afterwards.

That is one release, from my own work, and it does not prove one vendor reviews better than another. What it showed me is that a reviewer from outside the family was catching a different kind of problem: the kind that comes from reading the requirement differently.

There is a gradient here worth being precise about. Cross-vendor review in a separate process is how I work. Where a second vendor is not available, the fallback is a different model from the same vendor in the reviewer seat. I use that too and it helps, but less, and I would not describe the two as equivalent.

Beyond code

I expected this to be a code review practice. It has turned out to apply to anything a model produces where being right matters more than sounding right.

Research is the clearest case. Claude Code gathers and synthesises, then Codex gets the findings and the original question and is asked to challenge them: which claims rest on a single source, which are asserted without support, where the conclusion goes further than the evidence. The first pass rarely flags its own weak points. The second pass is set up to look for nothing else.

That changes the orchestrator's job. It stops being the agent that does the work and becomes the one that runs a structured disagreement between two models and decides what survives it.

Parallel streams, honestly counted

The orchestration builds separate git worktrees so several streams of work can run at once, each with its own review. I work solo, so in practice I am directing a small team of AI engineers.

The honest count is smaller than the tidy version. On the release I examined most closely, the work was sliced seven ways, but only three streams were genuinely independent. The rest were sequential: the API change had to land before the admin screens could use it. Forcing parallel worktrees onto sequential work adds overhead and buys nothing.

So I state the real number of independent streams before starting, and only split the work when it is above one. Parallelism is where this approach is easiest to oversell, and an oversold process gets dismissed by the first person who tries it on work that does not suit it.

What I have not tested

Everything above runs on hosted models. I have not yet put a local model in the reviewer position.

The case for trying is that the independence comes from the setup rather than from raw capability. A reviewer that cannot see the implementer's reasoning, and was trained differently, might be good enough for most routine reviews, cheaper, and would keep the code on your own hardware. The case against is that a reviewer too weak to follow the code will return cannot tell on everything, which looks like caution and is really noise.

I do not know where that line falls, and I am not going to guess at it in print. It is the next thing I intend to run.

If you are working out where review should sit in an AI-assisted delivery process, I am happy to compare notes. No pitch attached.