Project report / GitHub evidence
Claw-Eval Model Card
这是可索引项目报告证据页:它保留 Claw-Eval Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。
Claw-Eval Model Card
| Field | Value |
|---|---|
| Repository | claw-eval/claw-eval |
| Category | 可信 Agent 评测 |
| Stars / forks snapshot | 606 / 52 |
| Language | Python |
| License | NOASSERTION |
| Raw capture | raw-github/claw-eval_claw-eval.md |
| Updated by | hourly public metadata update, 2026-05-24 |
1. Role in Self Evolve
用 human-verified tasks、rubrics 和 Pass^3 多次运行逻辑降低 lucky-run 偏差,适合作为可信 agent evaluation 样本。
2. Working Principle
300 tasks + 2,159 rubrics + Pass^3 + full-trajectory auditing。
3. Evidence Path
web GitHub page captured March 2026 evaluation logic and task split signals.
4. Teaching Use
Use this card to explain whether a repository is a runtime, benchmark, harness-evolution loop, memory/skill substrate, or resource index. The key reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.
5. Limits
未复现 leaderboard;许可证字段未在公开页面片段中确认。