Evidence

Code evolution claims are not one ranking.

Score gains are weak evidence unless failed candidates, evaluator details, retained artifacts, and transfer checks are visible. Separate the mutable object before comparing systems.

ModeMutable objectExamplesBoundary
Self-modifying coding agentAgent code, tool policy, patch loop, or review logic changes.DGM, SICA, A-EvolveExamples are not uniformly reported by every system; inspect per-system reports.
Algorithm discoveryCandidate programs, kernels, heuristics, or optimization code change.AlphaEvolve, FunSearch, OpenEvolveStrong only when executable evaluators reject invalid programs.
Agent architecture searchAgent topology, control flow, or Python-coded design changes.ADAS, Agent Symbolic LearningTransfer and safety gaps matter more than one public score.
Prompt/program optimizerPrompt, symbolic module, or optimizer state changes.OPRO, DSPy, GEPACan overfit validation sets or judges.
Reflection and repair loopFailures become reusable critiques, tests, memories, or retry strategies.Reflexion, ReVeal, EvoMACFalse tests and memory pollution are the first risks to inspect.
Evidence and limits

The matrix is for claim triage, not final ranking.

Code-evolution systems differ by mutable object and evaluator strength. Treat the examples as anchors for review; cite specific score claims only after checking the original paper, benchmark protocol, retained artifacts, and failed-candidate disclosure.