Migrating from rouge-score to rouge-score-rs
rouge-score is pure Python, so scoring thousands of summaries or model outputs, or scoring them on every evaluation step, takes noticeable time. rouge-score-rs is a Rust port with the same API and the same scores, bit for bit. Most projects migrate by changing one line.
1. Swap the dependency
- rouge-score
+ rouge-score-rsAdd extras if you use them: rouge-score-rs[aggregate] for scoring.BootstrapAggregator (NumPy) and rouge-score-rs[sentences] for split_summaries=True (NLTK). Without extras the package has no Python dependencies. Wheels cover Python 3.9 and newer on Linux, macOS and Windows.
2. Change the import
- from rouge_score import rouge_scorer, scoring
+ from rouge_score_rs import rouge_scorer, scoringThe rest stays: RougeScorer(rouge_types, use_stemmer=..., split_summaries=..., tokenizer=...), score, score_multi, scoring.BootstrapAggregator and tokenize.tokenize. Results are the same Score(precision, recall, fmeasure) tuples and compare equal to the ones from rouge-score.
3. Replace loops and process pools with score_batch
If you parallelised rouge-score with multiprocessing, one call now does the same on all cores, without starting processes or pickling data:
- with Pool() as pool:
- scores = pool.starmap(scorer.score, zip(references, predictions))
+ scores = scorer.score_batch(references, predictions)The result is a list of dicts, one per pair and in order, equal to what the loop returns. RAYON_NUM_THREADS limits the threads. Pairs are scored one by one, without the parallel speedup, when the scorer has a custom Python tokenizer or split_summaries=True, and in a process forked after a parallel call. Pools keep working too: scorers can be pickled.
4. Hugging Face Evaluate
- rouge = evaluate.load("rouge")
+ rouge = evaluate.load("SyntaxSpirits/rouge")The SyntaxSpirits/rouge metric accepts the same arguments (rouge_types, use_aggregator, use_stemmer, tokenizer) and needs pip install evaluate "rouge-score-rs[aggregate]>=0.1.1". Pass the same seed to evaluate.load as before if you compare aggregated scores: Evaluate fixes the bootstrap seed when the metric is loaded.
5. Check parity on your own data
Before switching a reporting pipeline, run both packages on a sample of your real data. Keep rouge-score installed for this step only:
from rouge_score import rouge_scorer as reference
from rouge_score_rs import rouge_scorer as fast
types = ["rouge1", "rouge2", "rougeL", "rougeLsum"]
old = reference.RougeScorer(types, use_stemmer=True)
new = fast.RougeScorer(types, use_stemmer=True)
mismatches = [
i
for i, (ref, pred) in enumerate(zip(references, predictions))
if old.score(ref, pred) != new.score(ref, pred)
]
print(f"{len(mismatches)} of {len(references)} pairs differ")The comparison is exact, with no tolerance. If any pair differs, please open an issue with the two texts and the options: that is a bug.
Known differences
- An invalid ROUGE type raises
ValueErrorwhen the scorer is created, not on the first call toscore. Names must be exactlyrouge1…rouge9,rougeLorrougeLsum. - Texts must be
str;rouge-scorealso accepts UTF-8bytes. - A custom tokenizer written in Python still runs in Python, so it limits the speedup and
score_batchscores such pairs one by one.
Non-English text
Like the original, the default tokenizer drops everything except ASCII letters and digits, so Cyrillic, Arabic or Chinese text scores zero. If you need ROUGE for those languages, use the opt-in UnicodeTokenizer or CharacterTokenizer described on the rougers page, and report the profile name, since these scores are not comparable with rouge-score defaults.
Rolling back
Revert the import and the dependency. Nothing is stored in a different format, and the scores you computed are the ones rouge-score would have given.
Ready to switch?
pip install rouge-score-rsSpeed numbers and method: benchmarks. Source: github.com/SyntaxSpirits/rougers.