Waza Agent Skill Evaluation CLI Model Card
这是可索引项目报告证据页:它保留 Waza Agent Skill Evaluation CLI Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。
Waza Agent Skill Evaluation CLI Model Card
| Field | Value |
|---|---|
| Repository | microsoft/waza |
| Category | Waza Agent Skill Evaluation CLI |
| Stars / forks snapshot | 904 / 49 |
| Language | Go |
| License | MIT |
| Raw capture | raw-github/microsoft_waza.md |
| Updated by | hourly public metadata update, 2026-05-26 01:38 +0800 |
1. Role in Self Evolve
Waza is Microsoft’s Go CLI / framework for agent skills: it scaffolds skills and eval suites, runs benchmark tasks, compares models, checks coverage, and turns SKILL.md assets into measurable quality gates.
2. Working Principle
SKILL.md asset -> eval scaffold -> benchmark run -> grader/coverage report -> skill quality gate
3. Evidence Path
web-observed GitHub page showed 709 commits, MIT license, Go CLI, 904 stars, 49 forks, and README language for creating, testing, measuring and improving AI agent skill quality with eval.yaml suites, graders, multi-model comparison, coverage grids and CI comments; shell GitHub API freshness was blocked.
4. Teaching Use
Use this card to explain why skill evolution needs measurable evaluation and coverage gates before prompt or SKILL.md files can be treated as reusable capability.
5. Limits
The repository was not cloned in this iteration; no local install, benchmark rerun, security review, skill execution or CI gate was performed. Counts and claims are visible public-page signals unless independently revalidated later.