ROUGE benchmarks on real summarization data
How long it takes to score 2,000 summaries from CNN/DailyMail, XSum and PubMed with rouge-score 0.1.2 and rouge-score-rs. The numbers apply to this sample, these package versions and this machine; the commands at the end reproduce them.
30–80×
faster on one core, plain score loop
83–353×
faster with score_batch on 8 threads
0 / 6,000
score mismatches, compared without tolerance
Method
- 2,000 rows sampled uniformly from each full test split (
random.Random(1729)), from pinned dataset revisions verified by SHA-256. - Predictions are lead sentences: 3 for CNN/DailyMail, 1 for XSum, 5 for PubMed. Sentences are joined with newlines so
rougeLsumsees them. - Each configuration runs in a fresh Python process. Timings are the median of three complete calls after one warmup; imports, input construction and parity checks are outside the timings, Python result construction is inside.
- Every precision, recall and F-measure from
rouge-score-rsmust equalrouge-scoreexactly, otherwise the run fails. - Apple M1 Pro, 8 logical CPUs, macOS 14.5, Python 3.12.11,
rouge-score-rs0.2.1 (Git revision8e7aa7d) built in release mode, run on 2026-10-11.
| Dataset | Article words | Reference words | Prediction words |
|---|---|---|---|
| CNN/DailyMail | 687.5 | 54.7 | 75.8 |
| XSum | 381.8 | 20.9 | 19.9 |
| PubMed | 3038.3 | 205.0 | 150.4 |
Four ROUGE types
rouge1, rouge2, rougeL and rougeLsum; seconds for 2,000 pairs. The batch columns are score_batch with 1, 2, 4 and 8 threads.
| Dataset | Stem | rouge-score loop | rs loop | batch 1 | batch 2 | batch 4 | batch 8 | Speedup, batch 8 |
|---|---|---|---|---|---|---|---|---|
| CNN/DailyMail | no | 2.5209 | 0.0428 | 0.0442 | 0.0257 | 0.0245 | 0.0105 | 239.0× |
| CNN/DailyMail | yes | 5.2720 | 0.0929 | 0.0851 | 0.0462 | 0.0251 | 0.0175 | 300.4× |
| XSum | no | 0.3758 | 0.0126 | 0.0103 | 0.0066 | 0.0045 | 0.0045 | 82.6× |
| XSum | yes | 1.2187 | 0.0247 | 0.0237 | 0.0132 | 0.0079 | 0.0062 | 196.2× |
| PubMed | no | 13.6516 | 0.1693 | 0.1682 | 0.0867 | 0.0464 | 0.0438 | 311.4× |
| PubMed | yes | 19.6462 | 0.2804 | 0.3042 | 0.1479 | 0.0993 | 0.0557 | 352.5× |
Longer texts gain more: PubMed abstracts are about ten times longer than XSum summaries, and stemming is costly in Python.
Through Hugging Face Evaluate
evaluate.load("rouge") against evaluate.load("SyntaxSpirits/rouge"), both with seed=1729 and the default 1,000-sample bootstrap aggregation; aggregated results must match exactly. Seconds, Google / SyntaxSpirits.
| Dataset | Stem | Scoring | compute() | Speedup |
|---|---|---|---|---|
| CNN/DailyMail | no | 2.8886 / 0.0098 | 2.6896 / 0.2335 | 11.5× |
| CNN/DailyMail | yes | 5.1605 / 0.0489 | 5.3930 / 0.3006 | 17.9× |
| XSum | no | 0.3841 / 0.0039 | 0.7498 / 0.2258 | 3.3× |
| XSum | yes | 1.2187 / 0.0065 | 1.4617 / 0.2233 | 6.5× |
| PubMed | no | 13.2295 / 0.0437 | 21.5716 / 0.3191 | 67.6× |
| PubMed | yes | 20.0144 / 0.0579 | 20.6311 / 0.2868 | 71.9× |
The end-to-end gain is smaller than the scoring gain because bootstrap aggregation and input handling remain in Python, at about 0.2–0.3 s per call here.
Other ROUGE packages
rouge-rust computes rouge1, rouge2 and rougeL without stemming. On that subset it matched rouge-score on all 6,000 pairs; rouge-score-rs was 1.3–4.6× faster in a loop and 1.4–3.3× faster with 8 threads:
| Dataset | rouge-score | rs loop | rs batch 8 | rouge-rust loop | rouge-rust batch 8 |
|---|---|---|---|---|---|
| CNN/DailyMail | 1.3252 | 0.0172 | 0.0047 | 0.0384 | 0.0074 |
| XSum | 0.2082 | 0.0075 | 0.0029 | 0.0098 | 0.0042 |
| PubMed | 6.9406 | 0.0355 | 0.0081 | 0.1619 | 0.0269 |
Beyond that subset, rouge-score-rs also covers stemming, rougeLsum, score_multi, bootstrap aggregation, custom tokenizers and the Unicode profiles, with the rouge-score API. The rouge package (pltrdy) computes different scores, so it is not compared.
Memory
Whole-process peak RSS of each benchmark worker, including imports and results: rouge-score 74–96 MiB, rouge-score-rs 39–69 MiB.
Reproduce
From a checkout of SyntaxSpirits/rougers, with Python 3.12, Rust and uv:
uv venv .venv --python 3.12
uv pip install --python .venv/bin/python -r bindings/python/benchmarks/requirements.txt
source .venv/bin/activate
maturin develop --release -m bindings/python/Cargo.toml
python bindings/python/benchmarks/compare.py \
--pairs 2000 --seed 1729 --repeats 3 --output benchmark-resultsThe script downloads about 105 MB of pinned test data, prints Markdown tables and writes all 72 measurements to JSON. The full results and raw samples are in BENCHMARKS.md.