An autonomous tester can spend a fixed budget doing the same kind of work again and again. It may revisit a URL class, choose a familiar tool family, and receive another clean or error outcome without learning anything useful. My paper, The Fly That Stopped, studies one limited answer: make repetition itself visible to the scheduler, even when there is no scalar reward to optimize.
From a fruit fly to a scheduling prior
The mechanism is inspired by novelty processing in a fruit fly's mushroom body. It encodes a sparse view of the current state and keeps decaying habituation counters over structural URL classes and tool families. Repeated clean or error outcomes push the scheduler away from revisiting the same pattern. It is a reward-free scheduling prior, not a trained biological brain or a claim that a fruit fly has learned to pentest.
The earlier matched campaigns were useful partly because they exposed reward-accounting mistakes and tool-failure loops. In those campaigns, reward-driven components did not improve the tested primary outcomes over the reward-free mushroom-body condition. The paper therefore narrows the question instead of presenting a broad "smarter agent" claim.
What the confirmatory runs showed
After a pre-registered pilot and two confirmatory stages, eight of ten screened lab targets remained measurable in the second confirmatory stage; two error-heavy slow-XSS cases were excluded under the stated rules. On the measurable budget-hold population, the habituation-enabled scheduler reduced duplicate action ratios in all six non-tied target pairs, with two ties. The largest reported reduction was from 51 to 18 duplicate steps within a 60-step budget.
That is evidence for the complete scheduler on those measurable runs. It is not an isolated habituation ablation and not evidence of more vulnerabilities found. Duplicate discipline frees room in a budget; whether the agent uses that room to produce better, verified findings is a different experiment.
A second, more difficult question
The paper also examines a much larger circuit derived from the MaleCNS fly connectome. Five tested local-plasticity approaches did not produce action selectivity under the fixed readout. Readout plasticity produced qualified positive results on synthetic tasks, without establishing an advantage from the biological topology. That negative result belongs in the story: borrowing a biological motif for a simple prior is not the same as proving the larger neural circuit adds value.
The useful field takeaway is modest and testable: if an agent's actions are budgeted, record repeated selections and their outcomes before adding a reward learner. Read the full preprint for the preregistration, exclusions, population bounds, and input-integrity dependencies. The handbook's optional controller lessons provide a separate teaching path; they do not turn these lab results into a production claim.