Benchmark 表现
主指数只纳入 agent/code/web/app/algorithm-discovery 等自进化相关 benchmark family;AI-for-science adjacent evidence 单独展示,不进入主分。
Evolve-AGI Index 不是 AGI 能力分,而是一个 exploratory evidence-maturity worksheet,用来检查 AI Agent 自进化证据是否足够成熟。
The worksheet asks whether evidence is strong enough for the next review: can the benchmark be rerun, is the loop retained, are transfer and governance claims stated, and is the source chain visible?
Weights are editorial/proposed survey weights. This page does not claim peer-reviewed field consensus, confidence intervals, reviewer agreement, or sensitivity analysis. Treat it as a structured way to find weak evidence before making stronger claims.
主指数只纳入 agent/code/web/app/algorithm-discovery 等自进化相关 benchmark family;AI-for-science adjacent evidence 单独展示,不进入主分。
Top 自进化系统是否真的有可变对象、反馈信号、选择机制和保留机制。
综合 D2 证据强度与项目报告覆盖率,防止只有口号没有 raw / report。
改进是否能跨任务、跨环境或跨时间切片迁移,并有验证门约束。
系统是否有开源实现、文档、成本效率和实际采用价值。
结合证据分诊队列与 Star 活跃度,提示近期研究和工程关注度,不能单独当作价值排名。
看安全、成本、时间戳置信度和可审计边界是否跟上能力增长。