Choose a question, then inspect the evidence.
Do not start with a directory. Start with three questions: what counts, what evidence is strong, and what should I ignore for now?
Ten questions before trusting a self-evolving agent claim.
What counts as self-evolution?
Do not start with the project name. Check the mutable object, feedback, verification, retention, and rollback.
How does improvement happen?
Separate specification, search, evaluation, reflection, and archive pressure before comparing systems.
Which score gains are strong evidence?
Score gains are weak unless evaluator details, failed candidates, retained artifacts, and transfer checks are visible.
Can failure become reusable experience?
Memory and skill libraries matter only when they are retained, audited, and used in later runs.
Can the team structure itself evolve?
Look for changing roles, handoffs, topology, shared state, lineage, and independent graders.
How do we avoid benchmark hype?
Benchmarks are useful selection pressure, but they can leak, overfit, or reward the wrong behavior.
What is actually growing in 2026?
Current growth needs complete stargazer-history coverage; total stars are only historical adoption priors.
What should I read first?
Value LSH is a heuristic triage queue, not a final value judgment.
What do users really need?
Product value comes from reliability, transparency, control, and cost, not autonomy rhetoric.
What has this repository actually analyzed?
Separate raw collection, processed reports, public pages, and trend snapshots before trusting counts.
Boundary: this page gives English summaries for the core reader path. Some linked data views and reports are still Chinese-first, but their evidence sources are preserved.