Code and data for Pruning Laws for Large Language Models (EMNLP 2026).
A pruned LLM's downstream performance is predictable from two numbers: its unpruned performance and how much of it you removed.
L(L₀, r) = L₀ · P₀ · (1 − r)^α
L₀ is the unpruned score, r the pruning ratio, α the decay exponent, and
P₀ a bias term. Taking logs makes it a two-parameter linear regression, so
fitting it needs a handful of evaluations rather than a full sweep — and the
fitted coefficients transfer to models and pruners that were never in the fit.
git clone https://github.com/parmanu-lcs2/pruning_laws.git
cd pruning-laws
pip install -e .Python 3.10+. Dependencies are numpy, pandas, scipy, statsmodels and matplotlib.
from pruning_laws import load_results, fit_law, predict
df = load_results()
# Fit one law on LLaMA-7B under width pruning.
sub = df[(df.model == "LLaMA-7B") & (df.method == "WidthPruning")]
law = fit_law(sub.Average, sub.Average_base, sub.compression_ratio)
print(law)
# L = L0 * 0.734 * (1-r)^0.381 [adj R2=0.70, F=19.6, n=9]
# What would 45% pruning cost a model that scores 0.71 unpruned?
predict([0.71], [0.45], law.alpha, law.p0) # array([0.4152...])Reusing published coefficients on a model you have not swept — the zero-shot case — needs no fitting at all:
predict([0.71], [0.45], alpha=0.35, p0=0.85) # task-level Average coefficients
# array([0.4896...])python scripts/fit_laws.py # Table 2 (a, b, c)
python scripts/run_transfer.py # Figure 4 and Table 3
python scripts/fit_runtime_laws.py # Appendix Table 5
python scripts/make_figures.py # all figures
pytest # asserts every published numberThe test suite pins the published coefficients, so a change that moves them
fails loudly. REPRODUCING.md records two places where the paper's own tables
use inconsistent conventions, and how to reproduce each.
pruning_laws/
law.py the law: fit, predict, one-shot calibration
data.py loading and preprocessing
fitting.py fits at task / method / model / method-model level
evaluation.py held-out protocols (ratio-wise, extrapolation, model, method)
transfer.py transfer across model families and to unseen pruners
runtime.py speedup laws
plots.py figures
tables.py CSV and LaTeX export
scripts/ command-line entry points
data/ evaluation results (see data/README.md)
tests/ reproduction tests
The law can be fitted at four levels, differing in what is pooled first:
| Level | Pooled over | Use when |
|---|---|---|
| task | models and pruners | you want one number per task |
| method–task | models | you have chosen a pruner |
| model–task | pruners | you have chosen a model |
| method–model–task | nothing | both are fixed |
from pruning_laws.fitting import fit_method_task_level
fit_method_task_level(df)[["Method", "Task", "alpha", "P0", "adj_R2"]]extrapolation_rmse is the paper's "Test Error": the law is fitted on ratios
below a cut and tested at or above it, for every cut, and the errors averaged.
It answers the practitioner's question — how far past the ratios you have
measured can you trust the curve — rather than the easier interpolation
question. ratiowise_cv_rmse (leave-one-ratio-out) and modelwise_cv_rmse
(leave-one-model-out) are also reported.
from pruning_laws.transfer import method_transfer_table
from pruning_laws import load_baseline_methods
method_transfer_table(df, load_baseline_methods())Zero-shot reuses (α, P₀) unchanged. One-shot keeps α and re-estimates P₀
from a single pruning run. One-shot is not uniformly better: it helps when the
calibration ratio is representative of the regime you are predicting and hurts
when it is not. In the paper's table it improves three of four unseen pruners
and makes SVD-LLM worse.
Depth-pruning scores are smoothed. Layer-collapse pruners such as LaCo
change their merge plan abruptly as r crosses a block boundary, so scores can
jump upward or stay exactly flat between adjacent ratios. The loader carries the
previous ratio forward across those discontinuities. This is required to
reproduce the published depth coefficients and it visibly changes them, so it is
an explicit step: load_results(smooth_depth=False) turns it off.
Speedup is hardware-conditional; performance is not. The performance law
describes the network. The speedup law describes the network and the kernel
running it: depth and width pruning yield dense smaller sub-networks that are
faster anywhere, whereas unstructured pruning needs N:M sparse kernel support to
realise anything, which is why its exponent is near zero here. Only five models
have measured runtimes; see data/README.md.
@inproceedings{sengupta2026pruning,
title = {Pruning Laws for Large Language Models},
author = {Sengupta, Ayan and Chaudhary, Siddhant and Chakraborty, Tanmoy},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}MIT. See LICENSE.