"You have been given the complete record of an unusual sequence of events. Determine what the record ultimately requires you to figure out. Then figure it out."
That is the entire instruction. Ten documents follow. Two of them are decoys — they sound relevant but carry no formal weight. Notice what is being asked before you try to answer.
The record never asks anything outright. But its own constraints pin down exactly one consistent world — and that world demands one final action. The question the model must discover is:
"Who is behind the pattern in the record, and what must the corrected record require to be done?"
You do not need to solve the whole logic world (that took a SAT solver to verify). Just the human part — the final decision. What is the correct final action?
Here is the hidden world — the unique assignment of all eight variables that satisfies every constraint in the record:
How the record pins this down (each line is a clue you just read, formalized):
The decoys, for the record: items (9) and (10) — "the harbor fee was never paid" — appear twice, sound conspiratorial, and constrain nothing. Removing them cannot change the answer (that is a structural guarantee, not an assumption).
The marginalia: K0 ("Shakespeare wrote Hamlet") is an essential knowledge anchor — the solver must connect it to date the events; K1 ("Penicillin was discovered by Fleming") is confirmatory — it corroborates but is not required.
Every EnigmaForge instance ships with machine-checkable certificates. For this one (seed 2026, world W002026):
Reproduce this instance, byte for byte, with one command:
python3 -m enigmaforge.pipeline --size small --seed 2026
Frontier models saw exactly what you saw in Step 1 — no question, no hint. The best recovered the world nearly perfectly; most could recover the facts yet failed to take the action the record required. That gap is what the benchmark measures.