Project report / GitHub evidence

GitTaskBench Model Card

这是可索引项目报告证据页:它保留 GitTaskBench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

GitTaskBench Model Card

FieldValue
RepositoryQuantaAlpha/GitTaskBench
CategoryRepo-Level Code Agent Benchmark Harness
Stars / forks snapshot255 / 20
LanguagePython
LicenseUnknown
Raw captureraw-github/quantaalpha_gittaskbench.md
Updated byhourly public metadata update, 2026-05-28 04:00 +0800

1. Role in Self Evolve

GitTaskBench is a repository-level benchmark for real-world coding-agent tasks from repository understanding through implementation and task delivery. It matters because self-evolving agents need repeatable harness control, measurable feedback loops, and reusable skill procedures before claiming stable improvement.

2. Working Principle

repo-level task suites -> environment setup and incremental bug-fixing traces -> cost-aware alpha metrics for code-agent performance -> multi-agent runner comparison across real repositories

3. Evidence Path

web-observed GitHub page showed 255 stars, 20 forks, 84 commits, README benchmark framing for repo-level code-agent evaluation, and no explicit repository license marker in the navigation snapshot. Shell GitHub API access remained blocked by DNS and local gh auth was invalid, so this card treats the snapshot as web-observed rather than API-verified.

4. Teaching Use

Use this card to explain Repo-Level Code Agent Benchmark Harness: it shows how harness/runtime/benchmark layers convert agent behavior into reproducible and auditable engineering workflows.

5. Limits

The repository was not cloned in this iteration; no benchmark run, plugin install, workflow execution, or agent loop experiment was executed. Counts and claims are visible public-page/search signals unless independently revalidated later.