Project report / GitHub evidence

WindowsAgentArena Model Card

这是可索引项目报告证据页:它保留 WindowsAgentArena Model Card 的源材料入口、机制线索和限制提醒;正文仍需 reader/editor 与 academic public-copy review 后才能当作最终结论引用。

WindowsAgentArena Model Card

FieldValue
Repositorymicrosoft/WindowsAgentArena
CategoryWindows OS Agent Benchmark
Stars / forks snapshot861 / 95
LanguagePython
LicenseMIT
Raw captureraw-github/microsoft_windowsagentarena.md
Updated byhourly public metadata update, 2026-05-24

1. Role in Self Evolve

Microsoft Windows Agent Arena 是面向多模态 OS agents 的 Windows 评测平台,使用 Docker、Windows 11 VM/golden image 和 OpenAI/Azure OpenAI endpoint 执行大规模 benchmark。

2. Working Principle

Windows 11 VM/golden image -> multimodal OS agent -> scalable benchmark execution -> report metrics

3. Evidence Path

web GitHub page observed technical report citation, Docker/WSL/Python prerequisites, Windows 11 VM golden image workflow, and local deployment commands; 861 stars, 95 forks, 86 commits; shell GitHub API was blocked by DNS and gh auth was invalid, so this run marks freshness as web-page observed rather than API verified.

4. Teaching Use

Use this card to explain whether a repository is a runtime, benchmark, harness-evolution loop, memory/skill substrate, or resource index. The key reading path is: raw capture -> classification row -> public site card -> project report -> aggregate GitHub analysis.

5. Limits

当前未克隆源码,未运行 benchmark 或 SDK examples;star/fork/commit 快照来自公开 GitHub 页面文本或可见页面片段。