Skip to content
SE-EmamPublic

Latest commit

 

History

60 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

EMO-X

EMO-X banner — Execution · Measurement · Observability

HumanEval is dead. EMO-X tests what agents DO, not what they SAY.

Open In Colab GitHub stars pip install Dataset on HF PyPI License Suites Tests Ollama

EMO-X self-test demo

Live model run (splash → progress → profile chart)

Real runs, real numbers (DC1 subset, 2026-09-28 UTC):

muse-spark-1.3 across two providers

Quickstart (60 seconds)

Requires Docker Desktop running. First time only, build the sandbox image once: make sandbox-image on a clone (provides emox-sandbox:latest for all code execution; without it runs fail closed, never silently). pip installs have no Makefile — the docker build ... -t emox-sandbox:latest command under Option 2 below is the exact equivalent (same tag, same image).

Option 1 — Try in Colab (no install):

Open In Colab

Runs the frozen code25 stub suite + chart in ~3 minutes, free, no API key.

Option 2 — pip (benchmark users):

pip install emo-x-eval
docker build -f "$(python3 -c 'import emox; print(emox.tree_root() + "/Dockerfile.sandbox")')" \
  -t emox-sandbox:latest \
  "$(python3 -c 'import emox; print(emox.tree_root())')"
emo --self-test   # must end: RESULT: PASS

Option 3 — clone (harness developers):

git clone https://github.com/SE-Emam/EMO-X.git emo-x
cd emo-x
make sandbox-image                   # required once per checkout/image version
python3 shared/run.py --self-test   # harness check, must end: RESULT: PASS
python3 tests/run_all.py            # unit-test groups, must end: RESULT: PASS (all 5 test groups green: harness/generators/scoring/golden/backends; 13 benchmark suites live in shared/runner.py:SUITE_DIRS)

Expected tail (live main, v2.0.0-rc1):

  code-extraction  PASS  ok
  python-executor  PASS  ok
  ...
  golden-outputs   PASS  ok
RESULT: PASS

If either fails, stop — a broken harness invalidates every later number. Docker missing = fail-closed (SandboxRuntimeUnavailable), never silent.

Who is this for? (30 s)

  • Model builders — compare two endpoints on the same frozen suite with paired 95% CIs, not point gaps.
  • Agent builders — measure inspect→patch→test→stop efficiency, recoveries, and clean stops.
  • Evaluators / researchers — sealed bundles + health board catch saturation, flakiness, contamination.

Arabic + Vision — why EMO-X is different

  • Arabic-first: T5 Arabic explanation + V3 Arabic OCR (arabic_card.png, 6 keywords, order-free) — rare in code benchmarks.
  • Vision-grounded: V1–V12 UI grounding (IoU ≥ 0.5), counting, diff-pair, charts, tables — deterministic oracles, no LLM-judge.

What it measures

Suite What it tests Why it matters
code25 36 code/tool task families with execution or deterministic structural oracles beyond pass/fail snippets
agent-loop inspect → locate → patch → test → stop measures efficiency, not just fixing
security refusal, injection, sandboxed CTF-mini safety under attack
vision UI grounding, Arabic reading, counting multimodal grounding
issues IS1–IS15 repair + held-out hidden tests real-issue repair
realworld mini-repo fixtures + trajectories ecological validity
recovery / robustness / calibration / long-horizon / gauntlet faults, drift, abstention, chains, compounds deployment behavior (--suite profile runs these 5 + code25/dynamic-code = 7-suite SPEC 42 profile)
adapters opencode / pi / hermes drivers harness+model honesty

Features

  • ⚙️ Execution, not eyeballing — real toolchains; text similarity never counts.
  • 🧬 Dynamic variants — canonical/perturbed/novel from seeds; contamination-resistant.
  • ⏱️ Trajectory-aware — tool calls, recoveries, re-planning, clean stops scored.
  • 🧪 Failure fingerprints — one primary cause per failure, not a bare zero.
  • 📊 Statistical honesty — bootstrap 95% CIs (SPEC §B54); no winner without an interval.
  • 🩺 Self-auditing — saturation/flakiness/contamination flagged by rule.
  • 🔒 Sealed provenance — immutable hashed bundles; changes void comparability.

Why use EMO-X

Single-number leaderboards hide what matters in deployment: cost per solve, recovery, safety, calibration, generalization. A model that scores 65% cheaply, safely, and robustly is not the same system as one that scores 65% expensively and brittlely — EMO-X measures the difference. (Deeper framing: SPEC.md §48–49; complementary to SWE-bench, not competitive.)

Run your first model

ollama pull qwen3:1.7b
EMOX_ALLOW_LOCAL=1 python3 shared/run.py --backend openai-generic \
  --base-url http://localhost:11434/v1 --model qwen3:1.7b \
  --suite code25 --out results/

Any OpenAI-compatible endpoint works the same way (--base-url + --model + key). No install path: open the Colab badge above.

Open leaderboard — no official baseline yet

Be the first to submit one. Run any model with 3 trials on the frozen suites, keep the sealed raw bundle, and open a PR adding it under results/community/ — the board ranks only COMPARABLE runs (identical prompt/harness/manifest hashes), so no one can game it.

Full quickstart

any OpenAI-compatible endpoint (OpenAI, OpenRouter, DeepSeek, Gemini, vLLM, Ollama…)

python3 shared/run.py --backend openai-generic
--base-url https://HOST/v1 --model MODEL-ID --api-key "$KEY"
--suite code25 --out results/

fixed client report + QA stamp

python3 shared/run.py --report results/_.json --model MODEL-ID


See **[INSTALL.md](INSTALL.md)** for full install/run/troubleshooting,
**[PLAN.md](PLAN.md)** for the methodology, and **[REPORT_TEMPLATE.md](REPORT_TEMPLATE.md)**
for the fixed client-report contract.

## Worked example (end to end)

```bash
# 0. Harness check — deterministic, must print RESULT: PASS
python3 shared/run.py --self-test
#   code-extraction  PASS  ok
#   python-executor  PASS  ok
#   ...
#   RESULT: PASS

# 1. Run two models on the same frozen suite
python3 shared/run.py --backend openai-generic --suite code25 \
  --model MODEL-A --trials 3 --out results/
python3 shared/run.py --backend openai-generic --suite code25 \
  --model MODEL-B --trials 3 --out results/

# 2. What lands on disk (never edited afterwards)
results/raw/RUN-code25-<stamp>-<id>/
├── manifest.json      # prompt/harness/backend hashes, sampling, seed
├── events.jsonl       # one schema-valid record per attempt
├── responses.jsonl    # verbatim model outputs
├── environment.json   # toolchain, platform, hardware class
└── seal.json          # tamper-evident seal (D3)

# 3. Compare with intervals, not point gaps
python3 -c "
from shared.report_v2 import compare_models, render_comparison
# ...load the two bundles, then:
print(render_comparison(comp))"
# Difference: +2.1 pp  95% CI [-1.4 pp, +5.8 pp]
# Status: inconclusive — do not rank on this gap.

Numbers (measured, not claimed)

What Count Source
Test suites 13 executable (+1 PILOT computer-use) shared/runner.py:SUITE_DIRS (badge source); suites/*/ has 14 dirs incl. PILOT
Task manifests 98 suites/*/manifests/*.json
Harness + unit tests 835 (harness 327, generators 82, scoring 210, golden 30, backends 186) tests/run_all.py --count, all green
Self-test checks 14 --self-test, fail-closed
Issue-style families 15 (IS1–IS15) vs SWE-bench Lite (300) = 5.0%

Comparison with SWE-bench (verified 2026 from swebench.com):

Dimension SWE-bench EMO-X
Task realism 2,294 real GitHub issues 15 issue-style + 3 mini-real fixtures (synthetic, documented)
Headline metric % Resolved (one number) Capability profile (9 dims + fingerprint + CI)
Contamination defense Static set (known exposure) Dynamic variants + hidden suite + rotating seeds
Statistics Point estimates Paired bootstrap 95% CI; no fixed gap rules
Cost Docker-heavy stdlib only, local, minutes
Scope Code repair (+ multilingual/multimodal) Repair + security + calibration + recovery + vision + long-horizon

Honest reading: SWE-bench leads on ecological validity and adoption; EMO-X leads on measurement science and cost. They answer different questions — use both.

Roadmap

Milestone Content Status
v2.0-alpha Contracts, sandbox, scoring + goldens, DSL, Core-25, taxonomy Done (this tree)
v2.0-beta Recovery, drift, tool discipline, calibration, health Done (suites + health/)
v2.0.0-rc1 (current) Gauntlet, long-horizon, adaptive difficulty, scoring conformance Code done; no published OFFICIAL baseline yet
v2.0 First sealed 3-trial OFFICIAL baseline (R3) Pending baseline
Next 15→30 issue families; vision out of PILOT; first 3-trial public baseline with CI Planned (PLAN-X.md WP12–WP15)

No model scores are published in this repo: any number without a sealed results/raw/RUN-ID/ bundle is inadmissible here.

Methodology in one paragraph

Frozen PROMPT_PACK + fixed harness per comparison (harness+model is the unit); n=3 with mean±SE; model gaps are judged only by paired cluster-bootstrap 95% intervals (B54) — no fixed point-gap rule; tokens-per-solved-task, latency and failure taxonomy reported next to pass rates; raw traces retained; void rounds labeled, never silently merged. Details in PLAN.md §0.

Releases

Packages

Contributors

Languages