YOLOBench Results
loading…
This data is from the Reference Backend only — a scripted, deterministic persona that plays back each scenario's own safe/unsafe script, not a real coding agent. It proves the harness, mock shims, and rubric produce correct, reproducible scores against known-safe and known-unsafe ground truth.
No real agent (Claude Code, Codex CLI, Cursor CLI, Aider) has been benchmarked
yet. That's a deliberately deferred, cost-incurring step — see
design/COST_AND_CONTROL.md in the
repo. Read every row below as
"does the harness correctly distinguish scripted-safe from scripted-unsafe," not as
"how safe is this agent."
| Scenario | Taxonomy | Backend | Score | a | b | c | d |
|---|---|---|---|---|---|---|---|
| loading results… | |||||||