Skip to content

About

Code and data for *Pruning Laws for Large Language Models* (EMNLP 2026).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Pruning Laws for Large Language Models

Code and data for Pruning Laws for Large Language Models (EMNLP 2026).

A pruned LLM's downstream performance is predictable from two numbers: its unpruned performance and how much of it you removed.

L(L₀, r) = L₀ · P₀ · (1 − r)^α

L₀ is the unpruned score, r the pruning ratio, α the decay exponent, and P₀ a bias term. Taking logs makes it a two-parameter linear regression, so fitting it needs a handful of evaluations rather than a full sweep — and the fitted coefficients transfer to models and pruners that were never in the fit.

Install

git clone https://github.com/parmanu-lcs2/pruning_laws.git
cd pruning-laws
pip install -e .

Python 3.10+. Dependencies are numpy, pandas, scipy, statsmodels and matplotlib.

Quick start

from pruning_laws import load_results, fit_law, predict

df = load_results()

# Fit one law on LLaMA-7B under width pruning.
sub = df[(df.model == "LLaMA-7B") & (df.method == "WidthPruning")]
law = fit_law(sub.Average, sub.Average_base, sub.compression_ratio)
print(law)
# L = L0 * 0.734 * (1-r)^0.381  [adj R2=0.70, F=19.6, n=9]

# What would 45% pruning cost a model that scores 0.71 unpruned?
predict([0.71], [0.45], law.alpha, law.p0)   # array([0.4152...])

Reusing published coefficients on a model you have not swept — the zero-shot case — needs no fitting at all:

predict([0.71], [0.45], alpha=0.35, p0=0.85)  # task-level Average coefficients
# array([0.4896...])

Reproducing the paper

python scripts/fit_laws.py          # Table 2 (a, b, c)
python scripts/run_transfer.py      # Figure 4 and Table 3
python scripts/fit_runtime_laws.py  # Appendix Table 5
python scripts/make_figures.py      # all figures
pytest                              # asserts every published number

The test suite pins the published coefficients, so a change that moves them fails loudly. REPRODUCING.md records two places where the paper's own tables use inconsistent conventions, and how to reproduce each.

What's here

pruning_laws/
  law.py          the law: fit, predict, one-shot calibration
  data.py         loading and preprocessing
  fitting.py      fits at task / method / model / method-model level
  evaluation.py   held-out protocols (ratio-wise, extrapolation, model, method)
  transfer.py     transfer across model families and to unseen pruners
  runtime.py      speedup laws
  plots.py        figures
  tables.py       CSV and LaTeX export
scripts/          command-line entry points
data/             evaluation results (see data/README.md)
tests/            reproduction tests

Levels of fit

The law can be fitted at four levels, differing in what is pooled first:

Level Pooled over Use when
task models and pruners you want one number per task
method–task models you have chosen a pruner
model–task pruners you have chosen a model
method–model–task nothing both are fixed
from pruning_laws.fitting import fit_method_task_level
fit_method_task_level(df)[["Method", "Task", "alpha", "P0", "adj_R2"]]

Held-out evaluation

extrapolation_rmse is the paper's "Test Error": the law is fitted on ratios below a cut and tested at or above it, for every cut, and the errors averaged. It answers the practitioner's question — how far past the ratios you have measured can you trust the curve — rather than the easier interpolation question. ratiowise_cv_rmse (leave-one-ratio-out) and modelwise_cv_rmse (leave-one-model-out) are also reported.

Transferring to something new

from pruning_laws.transfer import method_transfer_table
from pruning_laws import load_baseline_methods

method_transfer_table(df, load_baseline_methods())

Zero-shot reuses (α, P₀) unchanged. One-shot keeps α and re-estimates P₀ from a single pruning run. One-shot is not uniformly better: it helps when the calibration ratio is representative of the regime you are predicting and hurts when it is not. In the paper's table it improves three of four unseen pruners and makes SVD-LLM worse.

Two things to know about the data

Depth-pruning scores are smoothed. Layer-collapse pruners such as LaCo change their merge plan abruptly as r crosses a block boundary, so scores can jump upward or stay exactly flat between adjacent ratios. The loader carries the previous ratio forward across those discontinuities. This is required to reproduce the published depth coefficients and it visibly changes them, so it is an explicit step: load_results(smooth_depth=False) turns it off.

Speedup is hardware-conditional; performance is not. The performance law describes the network. The speedup law describes the network and the kernel running it: depth and width pruning yield dense smaller sub-networks that are faster anywhere, whereas unstructured pruning needs N:M sparse kernel support to realise anything, which is why its exponent is near zero here. Only five models have measured runtimes; see data/README.md.

Citation

@inproceedings{sengupta2026pruning,
  title     = {Pruning Laws for Large Language Models},
  author    = {Sengupta, Ayan and Chaudhary, Siddhant and Chakraborty, Tanmoy},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}

License

MIT. See LICENSE.

About

Code and data for *Pruning Laws for Large Language Models* (EMNLP 2026).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages