dev.review
Manifesto

Human eyes on robot code.

Where it came from

Three convictions and a question.

This did not start as a position. It started as four things that arrived at roughly the same time, and only one of them was hard.

There was too much to review

Not difficult work. Just more of it than there were hours, and every item still needed a person's name against it at the end. That is the ordinary version of the problem and it is the one that started this.

Agents can do the reading

Hand one a diff and it finds the things worth finding. It opens every file, every time, and it is no less careful on the fourteenth pull request of the day than the first. The drafting was solved before anything else was.

Then they posted

A comment would land on a pull request before I had read a word of it, under my name or under its own, and either way something had been said on my behalf that I had not agreed to. An agent that posts has settled something on a person's behalf, and it is the person who answers for it. Nothing else here is held as firmly.

Which leaves the hard part

Eyes have to land on something, and something is not everything. An agent will produce forty findings on a change that warranted two, and a person who is made to read all forty stops reading properly by the eighth. Getting the right things in front of a person is the only one of the four that did not arrive with its own answer.

The argument below works that last one out, as a chain: grant the first line, and the conclusion is not a preference, it is what is left.

The same thing, formally

  1. Premise

    Some systems will always require full human eyes on every change.

    This is a property of the systems, not a gap in the tooling. Payments, clinical records, infrastructure that cannot be rolled back: the requirement comes from what happens when the change is wrong, and nothing about a better generator changes that. So it does not expire.

  2. Therefore

    Seeing is required.

    Eyes cannot land on a change they cannot see. Those systems therefore require a way to see what is changing.

  3. Therefore

    Seeing must be effective.

    Attention is finite. Seeing more than you can act on is not seeing, it is noise. The way must let you look with attention proportional to the change, and act on what you find.

  4. Therefore

    It must be neutral.

    A tool that owns the agent, the storage and the host makes your seeing contingent on that vendor's choices. The way must work with any agent, any storage, any host.

  5. Conclusion

    The systems that require human eyes require a neutral, effective way to see what is changing.

    That is the product. The timelessness follows from the first premise rather than sitting beside it.

Why the problem does not expire

Every argument that this category is temporary rests on agents getting good enough that nobody has to look. That is an argument about drafting quality, and it answers the wrong premise.

The requirement to look is not a statement about how good the code is.

It is a statement about what an organisation is willing to be accountable for. A regulator, an auditor, a board or a customer contract asks who read it. "The model was confident" has never been an answer to that question and is not becoming one.

So the volume of change goes up, the share of it that a person drafts goes down, and the number of changes a person has to put their name to goes up with the volume. The problem gets larger as the tooling improves.

Positions, not features

The reviewer is the accountable party

Not the bot, not the vendor, not the pipeline. Nothing is posted until a person says so, and it is posted under their name. A tool that posts on its own behalf has quietly moved the accountability somewhere nobody agreed to.

A finding is opted in, never opted out

Every comment starts outside the review, and the reader puts it in one at a time. The other order - everything staged, and a reader who has to catch and drop what should not go - makes silence the same as agreement. Silence is not a decision, and nothing should be able to post on the strength of one.

Clean lenses are content

Every lens is shown, including the ones with nothing flagged, each with a line on what was actually checked. Hiding the clean ones makes a review look shorter and makes thoroughness unprovable. The difference between checked and clean, and not looked at, is the entire value of a review.

Evidence beats assertion

QA evidence shows the change was driven, not only read. A recording of the app being exercised is a fact. A claim in a summary is a claim.

A failing QA run: a recorded browser session confirming a shop's own order confirmation page, with the finding it evidences named beneath it

The run above is a real one: it is what proved the finding on the other side of it, not a mockup of what proof would look like.

Refusing a landlord

Your agents, your storage, your host. There is no adapter that points at us and no tier where we hold your files. You can leave, which is the reason to stay.

Use it ↓ Neutrality