Project report / GitHub evidence

kbench Model Card

这是可索引项目报告证据页:它保留 kbench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

kbench Model Card

FieldValue
RepositoryshareAI-lab/kbench
CategoryAgent Harness Benchmark CLI
Stars / forks snapshot10 / 1
LanguageTypeScript / Python
LicenseApache-2.0
Raw captureraw-github/shareai-lab_kbench.md
Updated byhourly public metadata update, 2026-05-26 02:39 +0800

1. Role in Self Evolve

kbench normalizes SWE, Terminal-Bench 2.0, tau-bench and Standardized Agent Exams through one CLI and harness contract, including Codex, Claude Code, Gemini CLI, kode-agent-sdk and custom adapter paths.

2. Working Principle

benchmark bridge -> kbench CLI -> built-in or custom agent harness -> standardized run artifacts

3. Evidence Path

web-observed GitHub page showed 10 stars, 1 fork, 12 commits, Apache-2.0 license, TypeScript/Python/Shell stack, support for SWE, TB2, Tau and SAE benchmarks, built-in harnesses for kode-agent-sdk, codex, claude-code and gemini-cli, generated custom adapters, and standardized run/trace artifacts; shell GitHub API freshness was blocked.

4. Teaching Use

Use this card to explain why self-evolving agents need a stable harness adapter contract before benchmark comparisons become trustworthy.

5. Limits

The repository was not cloned in this iteration; no local install, benchmark rerun, security review, skill execution or CI gate was performed. Counts and claims are visible public-page signals unless independently revalidated later.