Evolve-AGI / English mirror

Use it as a review worksheet, not as an AGI scoreboard.

Evolve-AGI Index 不是 AGI 能力分,而是一个 exploratory evidence-maturity worksheet,用来检查 AI Agent 自进化证据是否足够成熟。

The worksheet asks whether evidence is strong enough for the next review: can the benchmark be rerun, is the loop retained, are transfer and governance claims stated, and is the source chain visible?

Boundary

The score does not decide which system is best.

Weights are editorial/proposed survey weights. This page does not claim peer-reviewed field consensus, confidence intervals, reviewer agreement, or sensitivity analysis. Treat it as a structured way to find weak evidence before making stronger claims.

Signals

Seven questions behind the worksheet.

Weight 18%

Benchmark 表现

主指数只纳入 agent/code/web/app/algorithm-discovery 等自进化相关 benchmark family;AI-for-science adjacent evidence 单独展示,不进入主分。

score 80.1/100 paper-drafts/appendix.tex + site/src/data/survey.ts + research/ranking-framework/radar-profiles.json
Weight 20%

核心闭环强度

Top 自进化系统是否真的有可变对象、反馈信号、选择机制和保留机制。

score 80.2/100 site/src/data/rankings.ts + research/ranking-framework/README.md
Weight 18%

证据链可信度

综合 D2 证据强度与项目报告覆盖率,防止只有口号没有 raw / report。

score 78.7/100 analysis/github-project-data-analysis.json + projects/INDEX.md
Weight 14%

迁移与验证

改进是否能跨任务、跨环境或跨时间切片迁移,并有验证门约束。

score 65.2/100 paper-drafts/ch5-evaluation.tex + site/src/data/rankings.ts
Weight 12%

可运行与可复用

系统是否有开源实现、文档、成本效率和实际采用价值。

score 58.8/100 site/src/data/rankings.ts + site/src/data/starAnalysis.ts
Weight 10%

领域动量

结合证据分诊队列与 Star 活跃度,提示近期研究和工程关注度,不能单独当作价值排名。

score 68.9/100 site/src/data/analysis.json + site/src/data/starAnalysis.ts
Weight 8%

治理成熟度

看安全、成本、时间戳置信度和可审计边界是否跟上能力增长。

score 65.1/100 site/src/data/rankings.ts + output/raw-github-timestamp-index.md