Project report / GitHub evidence

evmbench Model Card

这是可索引项目报告证据页:它保留 evmbench Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

evmbench Model Card

FieldValue
Repositoryparadigmxyz/evmbench
CategorySmart Contract Agent Benchmark Harness
Stars / forks snapshot421 / 62
LanguageTypeScript / Python
LicenseApache-2.0
Raw captureraw-github/paradigmxyz_evmbench.md
Updated byhourly public metadata update, 2026-05-26 02:39 +0800

1. Role in Self Evolve

evmbench is a domain-specific benchmark and harness for LLM agents that find and exploit smart-contract bugs, wrapping Codex detect-mode workers, job queues, secret handling, result validation and a report UI.

2. Working Principle

contract upload -> sandboxed Codex detect worker -> JSON vulnerability report -> UI/report validation

3. Evidence Path

web-observed GitHub page showed 421 stars, 62 forks, 5 commits, Apache-2.0 license, TypeScript/Python/Docker stack, frontend/backend/worker architecture, pinned frontier-evals reference, Codex detect-mode worker flow, parseable JSON vulnerability validation, and explicit security guidance for untrusted code; shell GitHub API freshness was blocked.

4. Teaching Use

Use this card to show how domain benchmarks package sandboxing, secrets, validation and report surfaces around an agent instead of only scoring final text.

5. Limits

The repository was not cloned in this iteration; no local install, benchmark rerun, security review, skill execution or CI gate was performed. Counts and claims are visible public-page signals unless independently revalidated later.