An agent can produce a plausible finding with a real quote from a response and still be wrong about the vulnerability. A quotation proves that some text appeared. It does not prove that authorization failed, a protected resource was exposed, or the claimed impact is real. The question behind my preprint is narrower than "can AI pentest?": does a verifier-and-acceptance stage change which candidate findings the agent actually ships?
The experiment, not the promise
The paper reports an exploratory pilot, a pre-registered confirmatory ablation, and a separate factorial study across two deliberately vulnerable lab targets. In the ablation, the full stage pairs a model verifier with deterministic code that acts on its verdicts. Removing that combined stage changed pre-report suppression and lowered model-blinded shipped precision in the reported runs. The factorial study then switched the model verifier and acceptance rules independently; suppression was attributable to the model verifier in that setting, not to deterministic acceptance rules alone.
That distinction matters. "The code disposes" means the host enforces and records the decision; it does not mean a code rule can determine the truth of every security claim. Nor does a model's confident verdict replace a human reviewer.
What a useful rejection looks like
Consider a candidate that calls an HTTP 200 response an authorization bypass. The verifier can ask whether the captured body contains the supposedly protected data or merely a login page. The system should preserve the candidate, the evidence reference, the verdict, and the reason it was suppressed or retained. A suppressed claim is not automatically harmless: false rejection is also a security failure. This is why the paper measures what the stage keeps as well as what it filters.
The full design retained most model-adjudicated true candidates in the factorial study, but missed its pre-registered non-inferiority criterion. Recall against an incomplete frozen ground-truth list did not differ significantly, and equivalence was not established. Those are constraints on the conclusion, not footnotes to discard.
What remains open
Independent human blind adjudication of the retained packets is pending. The precision and sensitivity endpoints are therefore supporting evidence rather than final human-validated results. The paper also discloses six audit-trail failures, including one in the evaluation tooling. Its lab results do not establish superiority over other agents or generalize automatically to client environments.
The practical takeaway is about the reporting boundary: an autonomous system needs a way to challenge its own claims before they become deliverables, and to leave a reviewable trail when it does. Read the full paper for the methods, exact statistics, deviations, and limits. The handbook turns the control ideas into offline exercises; its current harness is a separate reference, not the historical experimental system.