kbench Model Card
这是可索引项目报告证据页:它保留 kbench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。
kbench Model Card
| Field | Value |
|---|---|
| Repository | shareAI-lab/kbench |
| Category | Agent Harness Benchmark CLI |
| Stars / forks snapshot | 10 / 1 |
| Language | TypeScript / Python |
| License | Apache-2.0 |
| Raw capture | raw-github/shareai-lab_kbench.md |
| Updated by | hourly public metadata update, 2026-05-26 02:39 +0800 |
1. Role in Self Evolve
kbench normalizes SWE, Terminal-Bench 2.0, tau-bench and Standardized Agent Exams through one CLI and harness contract, including Codex, Claude Code, Gemini CLI, kode-agent-sdk and custom adapter paths.
2. Working Principle
benchmark bridge -> kbench CLI -> built-in or custom agent harness -> standardized run artifacts
3. Evidence Path
web-observed GitHub page showed 10 stars, 1 fork, 12 commits, Apache-2.0 license, TypeScript/Python/Shell stack, support for SWE, TB2, Tau and SAE benchmarks, built-in harnesses for kode-agent-sdk, codex, claude-code and gemini-cli, generated custom adapters, and standardized run/trace artifacts; shell GitHub API freshness was blocked.
4. Teaching Use
Use this card to explain why self-evolving agents need a stable harness adapter contract before benchmark comparisons become trustworthy.
5. Limits
The repository was not cloned in this iteration; no local install, benchmark rerun, security review, skill execution or CI gate was performed. Counts and claims are visible public-page signals unless independently revalidated later.