Memento 3 agent clears all 25 public ARC-AGI-3 games while keeping its language model frozen
Researchers at University College London and Huawei's Noah's Ark Lab report a mean RHAE score of 100.0, the benchmark's maximum. No weights are updated. The agent keeps a Markdown rulebook plus a compiled Python world model and rewrites both whenever its predictions fail. An ablation credits the rulebook with a 9% cut in actions and an 18% cut in agent turns, and a two-model run on one game dropped actions from 899 to 597. The paper, posted October 9, also includes an Atari Pong test in which a learned controller trained on 9,504 frames wins three games 21 to 0 without calling the model during play. The authors say that is 42 times more sample efficient than EfficientZero V2. ARC-AGI-3 was designed to expose how slowly AI learns new environments compared with people. If editable external memory can close that gap, startups can improve agents through scaffolding and memory design rather than paying for fine-tuning or reinforcement learning runs.