Project report / GitHub evidence

agent-skills-eval Model Card

这是可索引项目报告证据页:它保留 agent-skills-eval Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

agent-skills-eval Model Card

FieldValue
Repositorydarkrishabh/agent-skills-eval
CategoryAgent Skills Evaluation Harness
Stars / forks snapshot34 / 5
LanguageTypeScript
LicenseMIT
Raw captureraw-github/darkrishabh_agent-skills-eval.md
Updated byhourly public metadata update, 2026-05-25

1. Role in Self Evolve

agent-skills-eval 是面向 agent skills 的轻量评测 harness,用任务执行和结果检查把技能目录转成可比较的质量证据。

2. Working Principle

skill corpus -> task prompts -> execution/evaluation harness -> pass/fail evidence -> skill quality comparison

3. Evidence Path

web search result observed darkrishabh/agent-skills-eval as a public GitHub repository for evaluating agent skills; included because benchmark/eval coverage is a first-class user requirement. Shell GitHub API access remained blocked by DNS in this run, so this card treats the current snapshot as web-observed rather than API-verified.

4. Teaching Use

Use this card to explain Agent Skills Evaluation Harness in the raw -> classification -> project card -> site/report pipeline. The reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.

5. Limits

当前未克隆源码,未运行 benchmark、skill install flows、memory experiments、MCP servers、agent evolution loops、security scanners 或 production deployments;star/fork/commit/release 快照来自公开 GitHub 页面文本或可见页面片段。