agent-skills-eval Model Card
这是可索引项目报告证据页:它保留 agent-skills-eval Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。
agent-skills-eval Model Card
| Field | Value |
|---|---|
| Repository | darkrishabh/agent-skills-eval |
| Category | Agent Skills Evaluation Harness |
| Stars / forks snapshot | 34 / 5 |
| Language | TypeScript |
| License | MIT |
| Raw capture | raw-github/darkrishabh_agent-skills-eval.md |
| Updated by | hourly public metadata update, 2026-05-25 |
1. Role in Self Evolve
agent-skills-eval 是面向 agent skills 的轻量评测 harness,用任务执行和结果检查把技能目录转成可比较的质量证据。
2. Working Principle
skill corpus -> task prompts -> execution/evaluation harness -> pass/fail evidence -> skill quality comparison
3. Evidence Path
web search result observed darkrishabh/agent-skills-eval as a public GitHub repository for evaluating agent skills; included because benchmark/eval coverage is a first-class user requirement. Shell GitHub API access remained blocked by DNS in this run, so this card treats the current snapshot as web-observed rather than API-verified.
4. Teaching Use
Use this card to explain Agent Skills Evaluation Harness in the raw -> classification -> project card -> site/report pipeline. The reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.
5. Limits
当前未克隆源码,未运行 benchmark、skill install flows、memory experiments、MCP servers、agent evolution loops、security scanners 或 production deployments;star/fork/commit/release 快照来自公开 GitHub 页面文本或可见页面片段。