Skip to content

ROUGE benchmarks on real summarization data

How long it takes to score 2,000 summaries from CNN/DailyMail, XSum and PubMed with rouge-score 0.1.2 and rouge-score-rs. The numbers apply to this sample, these package versions and this machine; the commands at the end reproduce them.

30–80×

faster on one core, plain score loop

83–353×

faster with score_batch on 8 threads

0 / 6,000

score mismatches, compared without tolerance

Method

DatasetArticle wordsReference wordsPrediction words
CNN/DailyMail687.554.775.8
XSum381.820.919.9
PubMed3038.3205.0150.4

Four ROUGE types

rouge1, rouge2, rougeL and rougeLsum; seconds for 2,000 pairs. The batch columns are score_batch with 1, 2, 4 and 8 threads.

DatasetStemrouge-score looprs loopbatch 1batch 2batch 4batch 8Speedup, batch 8
CNN/DailyMailno2.52090.04280.04420.02570.02450.0105239.0×
CNN/DailyMailyes5.27200.09290.08510.04620.02510.0175300.4×
XSumno0.37580.01260.01030.00660.00450.004582.6×
XSumyes1.21870.02470.02370.01320.00790.0062196.2×
PubMedno13.65160.16930.16820.08670.04640.0438311.4×
PubMedyes19.64620.28040.30420.14790.09930.0557352.5×

Longer texts gain more: PubMed abstracts are about ten times longer than XSum summaries, and stemming is costly in Python.

Through Hugging Face Evaluate

evaluate.load("rouge") against evaluate.load("SyntaxSpirits/rouge"), both with seed=1729 and the default 1,000-sample bootstrap aggregation; aggregated results must match exactly. Seconds, Google / SyntaxSpirits.

DatasetStemScoringcompute()Speedup
CNN/DailyMailno2.8886 / 0.00982.6896 / 0.233511.5×
CNN/DailyMailyes5.1605 / 0.04895.3930 / 0.300617.9×
XSumno0.3841 / 0.00390.7498 / 0.22583.3×
XSumyes1.2187 / 0.00651.4617 / 0.22336.5×
PubMedno13.2295 / 0.043721.5716 / 0.319167.6×
PubMedyes20.0144 / 0.057920.6311 / 0.286871.9×

The end-to-end gain is smaller than the scoring gain because bootstrap aggregation and input handling remain in Python, at about 0.2–0.3 s per call here.

Other ROUGE packages

rouge-rust computes rouge1, rouge2 and rougeL without stemming. On that subset it matched rouge-score on all 6,000 pairs; rouge-score-rs was 1.3–4.6× faster in a loop and 1.4–3.3× faster with 8 threads:

Datasetrouge-scorers looprs batch 8rouge-rust looprouge-rust batch 8
CNN/DailyMail1.32520.01720.00470.03840.0074
XSum0.20820.00750.00290.00980.0042
PubMed6.94060.03550.00810.16190.0269

Beyond that subset, rouge-score-rs also covers stemming, rougeLsum, score_multi, bootstrap aggregation, custom tokenizers and the Unicode profiles, with the rouge-score API. The rouge package (pltrdy) computes different scores, so it is not compared.

Memory

Whole-process peak RSS of each benchmark worker, including imports and results: rouge-score 74–96 MiB, rouge-score-rs 39–69 MiB.

Reproduce

From a checkout of SyntaxSpirits/rougers, with Python 3.12, Rust and uv:

sh
uv venv .venv --python 3.12
uv pip install --python .venv/bin/python -r bindings/python/benchmarks/requirements.txt
source .venv/bin/activate
maturin develop --release -m bindings/python/Cargo.toml
python bindings/python/benchmarks/compare.py \
  --pairs 2000 --seed 1729 --repeats 3 --output benchmark-results

The script downloads about 105 MB of pinned test data, prints Markdown tables and writes all 72 measurements to JSON. The full results and raw samples are in BENCHMARKS.md.