Change one import
The scorer, the options and the returned Score tuples are the same as in rouge-score:
from rouge_score_rs import rouge_scorer # was: from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL", "rougeLsum"], use_stemmer=True)
scores = scorer.score(
"The quick brown fox jumps over the lazy dog",
"The quick brown dog jumps on the log.",
)
scores["rouge1"]
# Score(precision=0.75, recall=0.6666666666666666, fmeasure=0.7058823529411765)score_multi, custom tokenizers, split_summaries, scoring.BootstrapAggregator and tokenize.tokenize work as before, and scorers can be pickled for worker processes. The migration guide covers dependencies, multiprocessing and how to check parity on your own data.
Score many pairs on all cores
score_batch takes lists of references and predictions, releases the GIL and runs on every CPU core with the default tokenizer:
results = scorer.score_batch(references, predictions) # one dict per pair, in order
results[0]["rougeL"].fmeasureSet RAYON_NUM_THREADS to limit the threads. With a custom Python tokenizer or split_summaries=True, and in a process forked after a parallel call, pairs are scored one by one.
How much faster
2,000 pairs from each test set, lead sentences as predictions, rouge1, rouge2, rougeL and rougeLsum. Seconds, median of three runs on an Apple M1 Pro. Every precision, recall and F-measure matched rouge-score 0.1.2 exactly.
| Dataset | Stemming | rouge-score | rouge-score-rs | score_batch, 8 threads | Speedup |
|---|---|---|---|---|---|
| CNN/DailyMail | no | 2.521 | 0.043 | 0.011 | 239× |
| CNN/DailyMail | yes | 5.272 | 0.093 | 0.018 | 300× |
| XSum | no | 0.376 | 0.013 | 0.005 | 83× |
| XSum | yes | 1.219 | 0.025 | 0.006 | 196× |
| PubMed | no | 13.652 | 0.169 | 0.044 | 311× |
| PubMed | yes | 19.646 | 0.280 | 0.056 | 353× |
Through Hugging Face Evaluate, with its default bootstrap aggregation, compute is 3–72× faster, because aggregation and input handling stay in Python. The benchmark page has the method, memory use, a comparison with other ROUGE packages and the commands to reproduce it.
Other languages
The default tokenizer is the one from rouge-score: it keeps only ASCII letters and digits, so Ukrainian, Arabic, Hindi or Chinese text scores zero. Version 0.2 adds two opt-in tokenizers that run natively, also in score_batch:
from rouge_score_rs import rouge_scorer, tokenizers
default = rouge_scorer.RougeScorer(["rouge1"])
default.score("Уряд оголосив нові заходи.", "Нові заходи оголосив уряд.")["rouge1"]
# Score(precision=0.0, recall=0.0, fmeasure=0.0)
unicode = rouge_scorer.RougeScorer(["rouge1", "rougeL"], tokenizer=tokenizers.UnicodeTokenizer())
unicode.score("Уряд оголосив нові заходи.", "Нові заходи оголосив уряд.")["rouge1"]
# Score(precision=1.0, recall=1.0, fmeasure=1.0)
chinese = rouge_scorer.RougeScorer(["rouge1", "rougeL"], tokenizer=tokenizers.CharacterTokenizer())
chinese.score("今天天气很好", "今天天气不错")["rouge1"]
# Score(precision=0.6666666666666666, recall=0.6666666666666666, fmeasure=0.6666666666666666)UnicodeTokenizer(profileunicode-v1) splits words by the Unicode word boundary rules, for languages that separate words with spaces.CharacterTokenizer(profilecharacter-v1) gives character-level ROUGE, a common choice for Chinese.- Both normalise with NFKC and full case folding and use pinned Unicode 17.0 data, so tokens do not depend on the platform. Their scores are not comparable with
rouge-scoredefaults: report the profile name with your results.
Hugging Face Evaluate
The SyntaxSpirits/rouge metric takes the same inputs and options as Evaluate's built-in rouge and returns the same scores:
import evaluate
rouge = evaluate.load("SyntaxSpirits/rouge") # was: evaluate.load("rouge")
results = rouge.compute(predictions=predictions, references=references)It needs pip install evaluate "rouge-score-rs[aggregate]>=0.1.1". Optional backends for the built-in metrics are proposed upstream in huggingface/evaluate#827 and huggingface/lighteval#1426; both are open.
Rust
use rougers::{RougeScorer, RougeType};
let scorer = RougeScorer::new(["rouge1", "rougeL"], true)?; // true = Porter stemming
let scores = scorer.score(
"The quick brown fox jumps over the lazy dog",
"The quick brown dog jumps on the log.",
);
let rouge_l = scores.get(RougeType::L).unwrap();
println!("{:.4}", rouge_l.fmeasure); // 0.5882RougeScorer is Send + Sync, so one scorer can be shared across threads, for example with Rayon.
How parity is tested
The test suite fails on any difference in any float between rouge-score-rs and rouge-score 0.1.2:
- the stemmer against NLTK on all 235,976 words of a system dictionary;
- tokenization against Python's
str.lowerfor every Unicode code point; - thousands of random texts with repeated words, punctuation, digits and non-Latin scripts, plus property-based tests;
score_multi, custom tokenizers,split_summariesand bootstrap aggregation with a fixed NumPy seed.
Known differences: an invalid ROUGE type raises ValueError when the scorer is created rather than on the first call; texts must be str, not bytes; BootstrapAggregator needs the [aggregate] extra (NumPy) and split_summaries=True the [sentences] extra (NLTK). Without them the package has no Python dependencies.
Questions
Will my reported scores change?
No, with the default tokenizer. The scores are the same floats rouge-score 0.1.2 returns, so numbers in papers and dashboards stay comparable. Only the opt-in Unicode and character tokenizers give different scores.
Do I need a Rust toolchain?
No. PyPI has prebuilt abi3 wheels for Linux (x86_64 and aarch64, glibc and musl), macOS (Intel and Apple Silicon) and Windows x86_64, one wheel per platform for every Python from 3.9. Other platforms build from the source distribution and need Rust.
Is this an official Google project?
No. It is an independent port under the same Apache-2.0 licence as rouge-score and NLTK, whose scoring, tokenization and stemming logic it reproduces.