Evidence
Code evolution claims are not one ranking.
Score gains are weak evidence unless failed candidates, evaluator details, retained artifacts, and transfer checks are visible. Separate the mutable object before comparing systems.
| Mode | Mutable object | Examples | Boundary |
|---|---|---|---|
| Self-modifying coding agent | Agent code, tool policy, patch loop, or review logic changes. | DGM, SICA, A-Evolve | Examples are not uniformly reported by every system; inspect per-system reports. |
| Algorithm discovery | Candidate programs, kernels, heuristics, or optimization code change. | AlphaEvolve, FunSearch, OpenEvolve | Strong only when executable evaluators reject invalid programs. |
| Agent architecture search | Agent topology, control flow, or Python-coded design changes. | ADAS, Agent Symbolic Learning | Transfer and safety gaps matter more than one public score. |
| Prompt/program optimizer | Prompt, symbolic module, or optimizer state changes. | OPRO, DSPy, GEPA | Can overfit validation sets or judges. |
| Reflection and repair loop | Failures become reusable critiques, tests, memories, or retry strategies. | Reflexion, ReVeal, EvoMAC | False tests and memory pollution are the first risks to inspect. |
Evidence and limits
The matrix is for claim triage, not final ranking.
Code-evolution systems differ by mutable object and evaluator strength. Treat the examples as anchors for review; cite specific score claims only after checking the original paper, benchmark protocol, retained artifacts, and failed-candidate disclosure.