OpenHands Benchmarks Model Card
这是可索引项目报告证据页:它保留 OpenHands Benchmarks Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。
OpenHands Benchmarks Model Card
| Field | Value |
|---|---|
| Repository | OpenHands/benchmarks |
| Category | OpenHands Agent Evaluation Harness |
| Stars / forks snapshot | 85 / 62 |
| Language | Python |
| License | MIT |
| Raw capture | raw-github/openhands_benchmarks.md |
| Updated by | hourly public metadata update, 2026-05-25 |
1. Role in Self Evolve
OpenHands Benchmarks 是 OpenHands V1 的 evaluation harness,用标准化 pipeline 测试软件工程、通用推理、Commit0 和 workplace safety 等真实任务上的 agent 能力。
2. Working Principle
OpenHands agent -> benchmark adapter -> standardized evaluation pipeline -> migration to V1 Software Agent SDK
3. Evidence Path
web GitHub page observed 414 commits, MIT license, Python/Shell/Jinja stack, migration to OpenHands Software Agent SDK V1, active SWE-Bench/GAIA/Commit0/OpenAgentSafety benchmark table, 85 stars and 62 forks. Shell GitHub API access remained blocked by DNS and local gh auth was invalid, so this card treats the current snapshot as web-observed rather than API-verified.
4. Teaching Use
Use this card to explain OpenHands Agent Evaluation Harness in the raw -> classification -> project card -> site/report pipeline. The reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.
5. Limits
当前未克隆源码,未运行 benchmark、SDK examples、skill install flows、memory experiments 或 production deployments;star/fork/commit 快照来自公开 GitHub 页面文本或可见页面片段。