Code is disposable. The harness is permanent.
IRONLOOP is a verification-first software engineering system for AI-assisted development. Five layers of automated verification that turn code from a liability into a commodity.
You don't write code anymore. You build a harness around code. When the harness passes, the code is correct. When it doesn't, the code gets regenerated until it does.
| Layer | Name | What happens | Driver |
|---|---|---|---|
| 1 | SPEC | Contracts, types, interfaces, failure modes, red tests | Human |
| 2 | GEN | Code generation + compiler feedback loop, strict lints | Agent |
| 3 | TEST | TDD loop + mutation testing (proves tests can fail) | Agent |
| 4 | SIM | Deterministic simulation, property tests, fuzzing | CI |
| 5 | PENTEST | Deterministic scanners, then multi-model attack | CI |
More tokens burned on verification than on generation.
If your verification budget is smaller than your generation budget, you are doing it wrong.
For an internal CLI: 30% generation, 70% verification. For a distributed system: 10% generation, 90% verification.
The more critical the system, the heavier the harness.
The ratio is measured, not declared: every layer records the tokens it consumed, and the PR carries the resulting ratio. A harness that does not report its ratio is not an IRONLOOP harness.
AI code generation alone produces correct-looking code that fails at runtime. Teams that hand developers an AI tool and nothing else see no durable productivity gain: they took the speed without the harness.
IRONLOOP wraps every line of generated code in verification:
SPEC → GEN → TEST → SIM → PENTEST
↑ ↓ ↓ ↓ ↓
└─────── FEEDBACK LOOPS ──────────┘
- Layer 1 defines what "correct" means and turns every failure mode into a red test, before any code exists.
- Layer 2 lets the compiler reject bad code. The compiler is your
first reviewer. Strict lints stop the agent from appeasing the
compiler (
.clone(),.unwrap(),unsafe) instead of solving the problem. - Layer 3 turns the red tests green, then proves they can fail with mutation testing.
- Layer 4 proves your code works in hell, not just in a clean room: deterministic simulation with recorded seeds, so every failure replays.
- Layer 5 finds what attackers will find, before they do: deterministic scanners first, then several AI models attacking the code in parallel. A finding only counts if it reproduces.
Rust gives an AI agent the tightest feedback loop available. The compiler enforces four things no other mainstream language enforces at compile time:
- Exhaustive matching: every enum variant and every
Resultbranch must be handled. - Non-ignorable errors:
Resultis#[must_use]; an unhandled error is a warning, and warnings are errors here. - Aliasing and lifetimes: the borrow checker rejects data races and use-after-free before anything runs.
- No implicit null:
Optionreplaces the billion-dollar mistake.
You change an interface used in 40 files, the compiler lists all 40
errors, the agent fixes them one by one. Every cannot borrow as mutable
is free training data.
Other languages compile loosely, so the agent gets its first real signal at runtime, one layer too late. Python and JavaScript give the agent almost nothing at Layer 2.
IRONLOOP targets Rust, and only Rust. The reason is the loop: Rust's compiler is strict enough that most wrong code never reaches Layer 3, and stable enough (no breaking changes in a decade) that the training data is uniform. Brownfield code may be in any language; the rewrite target is always Rust. IRONLOOP does not provide Go, Python or JS gates and will not.
All 5 layers from the first commit. Spec and red tests first, then generation, then green, then simulation, then pentest. Never generate code before the spec is written. Never declare done before Layer 3 passes.
Progressive adoption. Add tests first (capture current behavior), extract specs by reverse-engineering, add simulation on hot paths, add pentest on exposed surfaces, then rewrite modules one by one using expand-contract: add new form → migrate callers in batches → delete old form.
# Install via skills.sh
npx skills add edouard-claude/ironloop
# Or install a specific skill
npx skills add edouard-claude/ironloop --skill ironloop
For Claude Code:
claude plugins install edouard-claude/ironloop
Or just drop SKILL.md into your agent's skills directory.
| Layer | Tools |
|---|---|
| 2 | cargo check, cargo clippy -D warnings (pedantic), cargo fmt, cargo deny |
| 3 | cargo test, cargo llvm-cov, cargo mutants, proptest |
| 4 | turmoil, madsim, loom, cargo fuzz, miri, kani; nightly: Toxiproxy / Chaos Mesh |
| 5 | cargo audit, cargo deny, cargo geiger, semgrep; then multi-model attack |
skills/engineering/ironloop/
├── SKILL.md ← Main entry point (YAML frontmatter + instructions)
├── greenfield.md ← New project workflow
├── brownfield.md ← Legacy migration workflow
├── triggers.md ← When to activate each layer
├── references/ ← Detailed docs loaded on demand
│ ├── 1-spec.md ← Specification layer
│ ├── 2-gen.md ← Generation + compiler loop + anti-appeasement lints
│ ├── 3-test.md ← TDD layer + mutation testing
│ ├── 4-sim.md ← Deterministic simulation + infra chaos
│ ├── 5-pentest.md ← Deterministic scanners + multi-model pentest
│ ├── cost.md ← Gate cost budgets (per PR vs nightly)
│ └── agent-budget.md ← Agent token ceilings + decision journal
└── assets/ ← Templates and resources
├── spec.md ← Project spec template
├── lints.toml ← Cargo [lints] block to paste into Cargo.toml
└── sim.yaml ← Simulation config
This repository runs a CI check on every push: frontmatter present,
every relative link resolves, SKILL.md stays under 500 lines. A
verification harness that does not verify itself has no standing.
Roadmap: a dogfooded example service (examples/) running Layers 2 to 5
in GitHub Actions and publishing the generation/verification token ratio
on each PR.
Distilled from 15 years of building and operating backend systems and cloud platforms, then refined through systematic R&D with AI coding agents. The five-layer structure and the Iron Law owe a debt to Quentin Adam (Clever Cloud): a constrained language makes a better agent, and verification must cost more than generation.
IRONLOOP follows the Agent Skills open specification:
- SKILL.md with required
name+descriptionYAML frontmatter - Progressive disclosure:
references/loaded on demand, core instructions under 500 lines assets/for templates (spec.md, lints.toml, sim.yaml)- LLM-agnostic: works with Claude Code, OpenAI Codex, Gemini CLI, Cursor, and any agent supporting the Agent Skills standard
MIT, © 2026 Edouard Claude
Built with 🦾 by edouard-claude