Project report / GitHub evidence

SkillsBench Model Card

这是可索引项目报告证据页:它保留 SkillsBench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

SkillsBench Model Card

FieldValue
Repositorybenchflow-ai/skillsbench
CategoryAgent Skills Benchmark Harness
Stars / forks snapshot1300 / 312
Commits / issues / PRs snapshot385 / 12 / 53
LanguagePython
LicenseApache-2.0
Latest visible commit date2026-06-03
Raw captureraw-github/benchflow-ai_skillsbench.md
Updated byhourly public metadata update, 2026-06-05 11:00 +0800

1. Role in Self Evolve

SkillsBench evaluates how well AI agents actually use reusable skills across specialized multi-step workflows under deterministic and gym-style benchmark settings. It matters because self-evolving agents need explicit runtime, memory, skill, and benchmark substrates before their improvement claims become trustworthy.

2. Working Principle

define skill-centric tasks -> run agent plus skill compositions -> score verifier outputs and task success -> compare per-task and per-skill behavior across models and runtimes

3. Evidence Path

web-observed GitHub repo page and commit history showed about 1.3k stars, 312 forks, 12 issues, 53 pull requests, 385 commits, Apache-2.0 license, explicit Hugging Face dataset and task taxonomy surfaces, and latest visible commits on 2026-06-03. This iteration keeps freshness honest: the snapshot comes from the public GitHub page observed on 2026-06-05, while shell GitHub API access remained blocked in this workspace.

4. Teaching Use

Use this card to explain Agent Skills Benchmark Harness: it shows how swarm runtimes, skill optimizers, benchmark suites, browser harnesses, and memory middleware fit into the broader self-evolving-agent pipeline.

5. Limits

The repository was not cloned in this iteration; no benchmark run, workflow execution, or agent loop experiment was executed. Counts and claims are visible public-page signals unless independently revalidated later.