Show your loop #326
Replies: 3 comments
|
Maintainer note: this is the pinned-style hub for first runs and demos. If you just tried Latest release: v1.6.0 |
What I ran
Outcome
Optional proof
|
|
We have open-sourced LoopArena, a benchmark for comparing models as runtime Controllers of coding-agent loops. The evaluated component is the Controller. Across Controller-model comparisons, the coding Worker and execution setup are held fixed, so the benchmark compares how different models guide the same Worker under common conditions. LoopArena evaluates this role at three scopes: execution-validated next-step decisions, repeated control over task slices, and complete software tasks.
Disclosure: I am one of the authors/maintainers. We would be interested in hearing how others evaluate the model or policy responsible for runtime loop control. |
Uh oh!
There was an error while loading. Please reload this page.
Show your loop
This is the permanent show-and-tell thread for Loop Engineering.
Post anything you’re proud of, stuck on, or surprised by — real runs preferred over theory.
Template (copy/paste)
Start here if you’re new
Demo: Loop Ready score climbing
Week-one culture: report-only (L1) before auto-fix (L2). See LOOP.md.
Also useful
Reply below — first loops, failures, and score screenshots all count.
All reactions