Skip to content

Repository files navigation

DecisionBench


The evaluation ecosystem for decision models

License CI

Installation

pip install git+https://github.com/Hanno-Labs/decision-bench.git
uv add git+https://github.com/Hanno-Labs/decision-bench.git

Example Usage

Smoke-test a supported decision model on the pinned benchmark. This example uses Bosun v3.1 0.6B, whose native decision-token readout is supported directly by run-hf.

hf download Hanno-Labs/bosun-v3.1-0.6b \
  --revision aaa9dd06d4d6501b33df61942472fed9284bc5e6 \
  --local-dir models/bosun-v3.1-0.6b

decision-bench run-hf task_specs/decisionbench-dev.toml \
  models/bosun-v3.1-0.6b results/bosun-v3.1-0.6b --smoke

Before running anything else, see the complete supported adapters and models. If your model is listed, use its runner; only add an adapter when its native decision readout is not already supported.

Overview

📈 Leaderboard Compare reviewed results and filter by task, family, domain, or primitive
🏃 Get Started Install DecisionBench and run the frozen suite
📋 Tasks and Views Understand the 23,900 rows, nine families, three primitives, and reasoning track
🤖 Models See supported adapters and models, or add a new native readout contract
📊 Results Load, inspect, and submit reproducible results
🧪 Evaluation Learn the metrics, artifacts, and comparability rules
🤝 Contributing Add models, tasks, benchmarks, and result records

Contribute

Report bugs and request features for any DecisionBench component in the central issue tracker. Send code changes to the repository that owns that component.

Choose the path that matches what you want to bring to DecisionBench:

Add a compatibility adapter so DecisionBench can evaluate a new decision model.

Run a supported model and submit its reviewed, reproducible scores.

Contribute one dataset-backed decision problem with labels, provenance, and tests.

Curate existing tasks into a named evaluation for a domain or purpose.

Citing

DecisionBench is under active development. Until the benchmark paper is published, cite the repository and the individual datasets listed in the task catalog. Machine-readable citation metadata lives in CITATION.cff.

About

Open benchmark runtime for document-grounded decision models

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages