Project report / GitHub evidence

OpenClaw ClawBench Model Card

这是可索引项目报告证据页:它保留 OpenClaw ClawBench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

OpenClaw ClawBench Model Card

FieldValue
Repositoryopenclaw/clawbench
CategoryAgent Harness 评测诊断
Stars / forks snapshot97 / 18
LanguagePython
LicenseMIT
Raw captureraw-github/openclaw_clawbench.md
Updated byhourly public metadata update, 2026-05-24

1. Role in Self Evolve

把模型能力、插件栈、harness 配置和执行轨迹一起纳入评价,避免只看最终输出。

2. Working Principle

trace scoring + reliability metrics + seed-noise/capability-signal decomposition + dynamical-systems failure regime diagnostics。

3. Evidence Path

web GitHub page captured Core v1 2026-04-20 release notes and repository surface metadata.

4. Teaching Use

Use this card to explain whether a repository is a runtime, benchmark, harness-evolution loop, memory/skill substrate, or resource index. The key reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.

5. Limits

当前未运行 1,080-run sweep,也没有本地 GitNexus symbol graph;仅作为公开 metadata 与方法信号。