| APIs | Documentation | Status |
|---|---|---|
| Python | scorio.readthedocs.io | |
| Julia | mohsenhariri.github.io/scorio/julia | |
| JavaScript/TypeScript | mohsenhariri.github.io/scorio/javascript |
-
September 2026 🤗: Our reasoning datasets Scorio Trace, Scorio Lite, Scorio Math, and Scorio GPQA are now available on Hugging Face. Explore the tutorial notebooks for evaluation, ranking, and answer aggregation.
-
August 2026: Preprint of our paper "Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility" is available.
-
July 2026 📦: Scorio is now available on npm. Read our blog post.
-
April 2026 🎉: Our ranking paper "Ranking Reasoning LLMs under Test-Time Scaling" has been accepted to ACL 2026 Main Conference!
-
January 2026 🎉: Our paper "Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation" has been accepted to ICLR 2026!
This repository contains three packages:
| Dataset | Attempts | Notebooks |
|---|---|---|
| Scorio Trace | 192,000 | Explore - Eval - Rank - Aggregate |
| Scorio Lite | 1,211,520 | Explore - Eval - Rank - Aggregate |
| Scorio Math | 59,520 | Explore - Eval - Rank - Aggregate |
| Scorio GPQA | 1,152,000 | Explore - Eval - Rank - Aggregate |
Scorio Lite contains the same attempts as Scorio Math and Scorio GPQA without the top-20 candidate distributions; its meta-* configurations omit token lists for smaller downloads. See the dataset guide for loading details and schemas.
# Install from PyPI
pip install scorio
# Install latest from GitHub
pip install "git+https://github.com/mohsenhariri/scorio.git"
# Install a specific tag
pip install "git+https://github.com/mohsenhariri/scorio.git@python-v0.2.3"
# Install from local repository
pip install -e .
import numpy as np
from scorio import eval
# Outcomes R: shape (M, N) with integer categories in {0, ..., C}
R = np.array([[0, 1, 2, 2, 1],
[1, 1, 0, 2, 2]])
# Rubric weights w: length C+1
# Here: 0=incorrect(0.0), 1=partial(0.5), 2=correct(1.0)
w = np.array([0.0, 0.5, 1.0])
# Optional prior outcomes R0: shape (M, D)
R0 = np.array([[0, 2],
[1, 2]])
# Bayesian evaluation with prior
mu, sigma = eval.bayes(R, w, R0)
print(f"μ = {mu:.6f}, σ = {sigma:.6f}")
# Expected: μ ≈ 0.575, σ ≈ 0.084275
# Bayesian evaluation without prior
mu2, sigma2 = eval.bayes(R, w)
print(f"μ = {mu2:.6f}, σ = {sigma2:.6f}")
# Expected: μ ≈ 0.5625, σ ≈ 0.091998
# Weighted average
accuracy, accuracy_sigma = eval.avg(R, w)
print(f"Average = {accuracy:.6f}, σ = {accuracy_sigma:.6f}")using Pkg
# From local development
Pkg.develop(path="./julia/Scorio.jl")
# Or from Julia General Registry
# Pkg.add("Scorio")using Scorio
# Outcomes R: shape (M, N) with integer categories in {0, ..., C}
R = [0 1 2 2 1;
1 1 0 2 2]
# Rubric weights w: length C+1
# Here: 0=incorrect(0.0), 1=partial(0.5), 2=correct(1.0)
w = [0.0, 0.5, 1.0]
# Optional prior outcomes R0: shape (M, D)
R0 = [0 2;
1 2]
# Bayesian evaluation with prior
mu, sigma = bayes(R, w, R0)
println("μ = $mu, σ = $sigma")
# Expected: μ ≈ 0.575, σ ≈ 0.084275
# Bayesian evaluation without prior
mu2, sigma2 = bayes(R, w)
println("μ = $mu2, σ = $sigma2")
# Expected: μ ≈ 0.5625, σ ≈ 0.091998
# Weighted average
accuracy, accuracy_sigma = avg(R, w)
println("Average = $accuracy, σ = $accuracy_sigma")Bayesian performance evaluation with uncertainty quantification using the Bayes@N framework.
R:M × Ninteger matrix with entries in{0, ..., C}(outcomes for M questions over N trials)w: lengthC+1float vector of rubric weights mapping categories to scoresR0(optional):M × Dinteger matrix of prior outcomes- Returns:
(mu, sigma)- posterior estimate and uncertainty
- Categories: Encode outcomes per trial as integers in
{0, ..., C} - Weights: Choose rubric weights
wof lengthC+1(e.g.,[0, 1]for binary outcomes) - Shapes:
RisM × N(M questions, N trials)R0isM × D(M questions, D prior trials)- Both must share the same
Mand category set
- Python 3.10+
- NumPy 2.0+
- Julia 1.6 or higher
- Node.js 18 or higher
If you use Scorio in your research, please cite the relevant papers:
@inproceedings{hariri2026dont,
title={Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation},
author={Hariri, Mohsen and Samandar, Amirhossein and Hinczewski, Michael and Chaudhary, Vipin},
booktitle={International Conference on Learning Representations},
year={2026},
url={https://proceedings.iclr.cc/paper_files/paper/2026/file/f04edfc65463d020629673a4bc4c58e7-Paper-Conference.pdf},
eprint={2510.04265},
archivePrefix={arXiv},
primaryClass={cs.AI},
note={Latest version available at \url{https://arxiv.org/abs/2510.04265}}
}@inproceedings{hariri2026ranking,
title={Ranking Reasoning {LLM}s under Test-Time Scaling},
author={Hariri, Mohsen and Hinczewski, Michael and Ma, Jing and Chaudhary, Vipin},
booktitle={Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)},
year={2026},
pages={33437--33478},
publisher={Association for Computational Linguistics},
doi={10.18653/v1/2026.acl-long.1544},
url={https://aclanthology.org/2026.acl-long.1544/},
eprint={2603.10960},
archivePrefix={arXiv},
primaryClass={cs.LG},
note={Latest version available at \url{https://arxiv.org/abs/2603.10960}}
}@misc{hariri2026testtime,
title = {Test-Time Scaling in Reasoning {LLM}s: Inference Regimes, Evaluation, and Reproducibility},
author = {Hariri, Mohsen and Chen, Weicong and Shahini, Nahal and Singh, Vikash and Ye, Kai and Samandar, Amirhossein and Ganguly, Debargha and Sankar, Sreehari and Zhang, Yanyan and Wang, Shouren and Peng, Jerry and Zhang, Biyao and Hinczewski, Michael and Chaudhary, Vipin},
year = {2026},
eprint = {2608.04001},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
doi = {10.48550/arXiv.2608.04001},
url = {https://arxiv.org/abs/2608.04001}
}Bug reports, feature requests, and new evaluation, ranking, or aggregation methods are welcome. See the contributing guide to get started.
|
@mohsenhariri |
@HarryHills3588 |
@NahalShahini989 |
@ben072292 |
|
@mecaneer23 |
@Amirsamandar |
@SreehariSankar |
@kv-248 |
View all contributors on GitHub.
This project is licensed under the MIT License - see the LICENSE file for details.
- Landing Page: mohsenhariri.github.io/scorio
- Python Docs: scorio.readthedocs.io
- Julia Docs: mohsenhariri.github.io/scorio/julia
- JavaScript/TypeScript Docs: mohsenhariri.github.io/scorio/javascript
- Repository: github.com/mohsenhariri/scorio
- Issues: github.com/mohsenhariri/scorio/issues
- Papers: