Skip to content
mohsenhaririPublic

About

Bayes@N [ICLR'26], Ranking LLMs [ACL'26 Main]: Statistical evaluation, comparison, and ranking of Large Language Models

Topics

Resources

Contributing

Stars

23 stars

Watchers

0 watching

Forks

Repository files navigation

Scorio

ICLR 2026 ACL 2026 arXiv: Test-Time Scaling License: MIT Python 3.10+ PyPI package Julia 1.6+ npm package Python Docs Julia Docs JavaScript/TypeScript Docs


Documentation

mohsenhariri.github.io/scorio

APIs Documentation Status
Python scorio.readthedocs.io ReadTheDocs
Julia mohsenhariri.github.io/scorio/julia GitHub Pages
JavaScript/TypeScript mohsenhariri.github.io/scorio/javascript GitHub Pages

News


Packages

This repository contains three packages:

  1. scorio - Python implementation PyPI
  2. Scorio.jl - Julia implementation Julia
  3. scorio - JS/TS implementation npm

Datasets

Dataset Attempts Notebooks
Scorio Trace 192,000 Explore - Eval - Rank - Aggregate
Scorio Lite 1,211,520 Explore - Eval - Rank - Aggregate
Scorio Math 59,520 Explore - Eval - Rank - Aggregate
Scorio GPQA 1,152,000 Explore - Eval - Rank - Aggregate

Scorio Lite contains the same attempts as Scorio Math and Scorio GPQA without the top-20 candidate distributions; its meta-* configurations omit token lists for smaller downloads. See the dataset guide for loading details and schemas.


Quick Start

Python (scorio)

Installation

# Install from PyPI
pip install scorio

# Install latest from GitHub
pip install "git+https://github.com/mohsenhariri/scorio.git"

# Install a specific tag
pip install "git+https://github.com/mohsenhariri/scorio.git@python-v0.2.3"

# Install from local repository
pip install -e .

Basic Usage

import numpy as np
from scorio import eval

# Outcomes R: shape (M, N) with integer categories in {0, ..., C}
R = np.array([[0, 1, 2, 2, 1],
              [1, 1, 0, 2, 2]])

# Rubric weights w: length C+1
# Here: 0=incorrect(0.0), 1=partial(0.5), 2=correct(1.0)
w = np.array([0.0, 0.5, 1.0])

# Optional prior outcomes R0: shape (M, D)
R0 = np.array([[0, 2],
               [1, 2]])

# Bayesian evaluation with prior
mu, sigma = eval.bayes(R, w, R0)
print(f"μ = {mu:.6f}, σ = {sigma:.6f}")
# Expected: μ ≈ 0.575, σ ≈ 0.084275

# Bayesian evaluation without prior
mu2, sigma2 = eval.bayes(R, w)
print(f"μ = {mu2:.6f}, σ = {sigma2:.6f}")
# Expected: μ ≈ 0.5625, σ ≈ 0.091998

# Weighted average
accuracy, accuracy_sigma = eval.avg(R, w)
print(f"Average = {accuracy:.6f}, σ = {accuracy_sigma:.6f}")

Julia (Scorio.jl)

Installation

using Pkg

# From local development
Pkg.develop(path="./julia/Scorio.jl")

# Or from Julia General Registry
# Pkg.add("Scorio")

Basic Usage

using Scorio

# Outcomes R: shape (M, N) with integer categories in {0, ..., C}
R = [0 1 2 2 1;
     1 1 0 2 2]

# Rubric weights w: length C+1
# Here: 0=incorrect(0.0), 1=partial(0.5), 2=correct(1.0)
w = [0.0, 0.5, 1.0]

# Optional prior outcomes R0: shape (M, D)
R0 = [0 2;
      1 2]

# Bayesian evaluation with prior
mu, sigma = bayes(R, w, R0)
println("μ = $mu, σ = $sigma")
# Expected: μ ≈ 0.575, σ ≈ 0.084275

# Bayesian evaluation without prior
mu2, sigma2 = bayes(R, w)
println("μ = $mu2, σ = $sigma2")
# Expected: μ ≈ 0.5625, σ ≈ 0.091998

# Weighted average
accuracy, accuracy_sigma = avg(R, w)
println("Average = $accuracy, σ = $accuracy_sigma")

Evaluation Functions

bayes(R, w, R0=None)

Bayesian performance evaluation with uncertainty quantification using the Bayes@N framework.

  • R: M × N integer matrix with entries in {0, ..., C} (outcomes for M questions over N trials)
  • w: length C+1 float vector of rubric weights mapping categories to scores
  • R0 (optional): M × D integer matrix of prior outcomes
  • Returns: (mu, sigma) - posterior estimate and uncertainty

Data and Shape Conventions

  • Categories: Encode outcomes per trial as integers in {0, ..., C}
  • Weights: Choose rubric weights w of length C+1 (e.g., [0, 1] for binary outcomes)
  • Shapes:
    • R is M × N (M questions, N trials)
    • R0 is M × D (M questions, D prior trials)
    • Both must share the same M and category set

Requirements

Python

  • Python 3.10+
  • NumPy 2.0+

Julia

  • Julia 1.6 or higher

JavaScript

  • Node.js 18 or higher

Citation

If you use Scorio in your research, please cite the relevant papers:

Bayesian Evaluation Framework

@inproceedings{hariri2026dont,
  title={Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation},
  author={Hariri, Mohsen and Samandar, Amirhossein and Hinczewski, Michael and Chaudhary, Vipin},
  booktitle={International Conference on Learning Representations},
  year={2026},
  url={https://proceedings.iclr.cc/paper_files/paper/2026/file/f04edfc65463d020629673a4bc4c58e7-Paper-Conference.pdf},
  eprint={2510.04265},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  note={Latest version available at \url{https://arxiv.org/abs/2510.04265}}
}

Ranking Methods

@inproceedings{hariri2026ranking,
  title={Ranking Reasoning {LLM}s under Test-Time Scaling},
  author={Hariri, Mohsen and Hinczewski, Michael and Ma, Jing and Chaudhary, Vipin},
  booktitle={Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)},
  year={2026},
  pages={33437--33478},
  publisher={Association for Computational Linguistics},
  doi={10.18653/v1/2026.acl-long.1544},
  url={https://aclanthology.org/2026.acl-long.1544/},
  eprint={2603.10960},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  note={Latest version available at \url{https://arxiv.org/abs/2603.10960}}
}

Aggregation Methods

@misc{hariri2026testtime,
  title         = {Test-Time Scaling in Reasoning {LLM}s: Inference Regimes, Evaluation, and Reproducibility},
  author        = {Hariri, Mohsen and Chen, Weicong and Shahini, Nahal and Singh, Vikash and Ye, Kai and Samandar, Amirhossein and Ganguly, Debargha and Sankar, Sreehari and Zhang, Yanyan and Wang, Shouren and Peng, Jerry and Zhang, Biyao and Hinczewski, Michael and Chaudhary, Vipin},
  year          = {2026},
  eprint        = {2608.04001},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2608.04001},
  url           = {https://arxiv.org/abs/2608.04001}
}

Contributing

Bug reports, feature requests, and new evaluation, ranking, or aggregation methods are welcome. See the contributing guide to get started.

Contributors


@mohsenhariri

@HarryHills3588

@NahalShahini989

@ben072292

@mecaneer23

@Amirsamandar

@SreehariSankar

@kv-248

View all contributors on GitHub.

Guidelines for coding agents

See AGENTS.md and CLAUDE.md for repository instructions.

License

This project is licensed under the MIT License - see the LICENSE file for details.


Links

About

Bayes@N [ICLR'26], Ranking LLMs [ACL'26 Main]: Statistical evaluation, comparison, and ranking of Large Language Models

Topics

Resources

Contributing

Stars

23 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages