Perslis Defense
PERSLIS DEFENSE · CHAPTER 8 · THE FAIL-FIRST MODEL

It fails first. It doesn’t fail the same way twice.

Every new system screws up. Two things matter: does it keep making the same mistake, and can a bad week teach it to ignore its orders? VDSG — our system for defense, built on Peel, the Fail-First Model — answers no to both. It learns from every failure. It cannot learn its way around its limits. Autonomy that accumulates operational experience.

What that bought in the games: a tank brain pruned of what it never used and what hurt it scored 48,588 against the native rules’ 29,275 over 80 fresh seeds (BattleZone). Learning on against learning off: 15 deaths an hour against 80 (Fallout). No neural network, no weights, no training run, no network link. What it knows is written on Peel cards. What it has learned is a table of counts you can read. Every refusal cites the failures that earned it.

Play ▸   The scoreboard →   What is a fail-first model? The math →   Read the paper ↗

VDSG playing DOOM E1M1 from the commanded console
VDSG playing DOOM (1993), E1M1, from the recorded console. Rules decide, experience learns, orders only narrow. No neural network in the loop.

What “fail-first” means for you

A fail-first system expects to fail, and learns only from failures it actually had. What makes that safe is an old engineering rule: when it fails, it must fail toward a controlled state, not an uncontrolled one.

Learning can change how it fights. It cannot change its limits.
In the words our tests use: it may rewrite what it believes works. It may never rewrite what it is authorised to do.

That is how it is built. It is not a safety certification, and nothing here is qualified. See Limits.

The loop, in plain terms

FAIL  →  OBSERVE  →  EXPLAIN  →  BUILD RULE  →  VERIFY  →  RETRY
  1. Fail. It dies, or takes a hit it did not need to take. That gets written down, not averaged away.
  2. Observe. It reads the situation at every decision from the engine’s own state: health, who is there, whether a fight or a conversation is going on. It does not guess from pixels what the game already states.
  3. Explain. It pins the failure on the decision that caused it. That is harder than it sounds. If a fight starts a moment after a conversation, the blame goes to the line that started it, not to the heal it tried in the middle of the fight. We got this wrong first (below), and it produced a pilot that died the same way forty-nine times.
  4. Build rule. A pattern gets banned only when the harm is clear. Either it got the pilot killed every time it was tried, at least twice. Or, after at least four tries, even the most generous estimate of its death rate (the lower edge of a 90% Wilson interval) is still worse than the pilot’s normal death rate plus a margin. One unlucky death is not a lesson. Every rule names the pattern and the experiences behind it.
  5. Verify. The rule works only inside the limits set before it. It can take a goal away, or reorder what is left. Nothing else. That boundary is stated as a proposition in the paper and pinned by tests.
  6. Retry. Next time it is in that spot, the banned move is off the table and everything else stays open. If experience bans every option, the least-bad one comes back — because standing still and dying is not adapting.

The chain of authority it cannot touch

Every decision passes through the same chain of authority. Learning goes last:

what the situation offers      rules.applicable(s)
  ∩ what the manual allows       Peel cards in force
  ∩ your standing orders         "don't fire", "hold position"
  = what it may do right now     ── authority ends here ──
  → experience narrows / reorders inside it     (the learner)
  → the rules pick one goal → the actuator presses the buttons

Three tests pin that separation. They run with the rest of the suite (219 passing as of 2026-09-26) on every change:

test_the_learner_can_never_add_a_goal
    however much experience it has, it cannot invent permission
test_the_learner_cannot_overrule_a_standing_order
    an order that narrows the set to FIGHT stays FIGHT, even after 20 deaths
test_ranking_is_a_permutation_and_nothing_more
    the learner may reorder the admissible goals, never add one

The second test is a deliberate choice: a human order outranks experience, even when experience says the order is getting the pilot killed. It can report that the order is costly. It cannot countermand it.

This is the part that matters in the fight. You and your chain of command set the authority, the orders and the ROE. The floor is the deterministic part that holds the machine to machine-readable limits derived from authorized policy, mission rules, safety limits and applicable ROE — and nothing it learns can loosen them. DoD Directive 3000.09 requires autonomous and semi-autonomous weapon systems to be designed so commanders and operators can exercise “appropriate levels of human judgment over the use of force.” We are building toward that framework. We do not claim to implement it or comply with it, and nothing here has left simulation.

Where its permissions come from

VDSG does not decide what it is allowed to do from its own experience. Permission comes from a document. In the Fallout (1997) test, the game’s printed survival manual is compiled into 2,250 Peel cards. The cards in force for the current situation allow the goals the pilot may pursue, and each permission cites its page. Peel is the same card format behind the science runtime: every fact is typed, sourced, and either complete or absent.

So there are two stores that never mix. The cards say what is allowed, with the source. The evidence table says what has gone wrong, with the experiences. Learning writes only to the evidence table. The card store is opened read-only, so nothing the pilot lives through can change what it is allowed to do.

The results, bad news first

Measured in-house, in games. Compare each row against its own baseline.
testresultwhat it means
Freeway
same failure memory as Invaders, no per-game strategy
9.2 vs 10.4
−12% · 3 rules, every one blocking up
Learning from failure hurts when the only way to score is also the dangerous move. It correctly learned that up is dangerous, and stopped scoring.
DOOM
rules, then rules + memory, 32 seeds
17.9 → 20.3
paired t = 0.78
Even, not a win. The rules beat random play (3.2) decisively; the memory on top does not separate from them.
Space Invaders
raw pixels, zero emulator RAM
200.6 vs 152.2
+32% · peak 276.2 @50 episodes
Learned refusal helps when a failure ends the game.
Fallout (1997)
scripted Shady Sands, driven through the real decision loop
dies once
then chooses the peaceful line
Before the blame fix, one live run died 49 times going back to the same guard. With no safe line available, it dies twice, then stops talking to that guard but keeps talking to the child.

The Fallout row is a scripted scene run through the real Session.tick(), not a long live campaign. In a longer live test since, learning on against learning off, same save and seed, 12 minutes each: 15 deaths an hour against 80. The Invaders and Freeway pair is the finding that matters most, and Evidence explains why the same code helps in one game and hurts in the other.

The worst failure we found was ours

The Fallout pilot kept walking back to the same guard, saying the same line, and dying. The learning worked. It was learning from the wrong moment. The guard ends the conversation first and draws a moment later, so the result of a line was read one tick too early. “Prepare to meet your maker”, said four times and followed by four fatal fights, was scored as fight 0%. Meanwhile the death memory blamed only the last six decisions, all inside the fight: 26 deaths were charged to HEAL, and the 446 replies that started those fights were charged nothing.

Blame assignment is where a learn-from-failure system goes wrong without telling you. Each of these bugs produced believable output until it was caught. This is the second one we have published; the first was the Atari death-animation defect.

Is any of this new?

Mostly no. Here is where each piece comes from:

What we claim, and only this: we know of no earlier system that combines all three of (1) rules built from failures it actually had, each citing those failures, (2) limits set before the learning, from a licensing document and human orders, that the learning provably cannot widen, and (3) no neural network anywhere in the loop. If you know of one, tell us and we will cite it here.

What we are not claiming

Play ▸   Talk to us ↗

Everything Perslis