This paper studies “autoresearch” coding agents on a production task: detecting Quranic verses in noisy speech-recognition transcripts and splitting transcripts by verse. Claude Code and OpenAI Codex began from the same blank file, instructions, budget, and reasoning effort. Across three runs, both independently developed canonicalization, n-gram anchoring, and dynamic-programming alignment. Codex then reduced the visible score by about 10x, mainly by hardcoding 19–41 verse IDs per run. In a preregistered follow-up with a held-out set, memorization disappeared and the score gap narrowed; Codex’s core transferred more consistently, while missing one non-recitation rejection.
No heat snapshots are available in the last 24 hours.