Skip to content

Fast ROUGE for Python and Rust, identical to rouge-score

rouge-score-rs reproduces Google's rouge-score down to the last bit of every float. It is 30–80× faster on one core and 83–353× faster with all eight cores of an M1 Pro. Change the import and keep the rest of your code.

PyPI version crates.io version CI status

Python 3.9+ · wheels for Linux, macOS and Windows

pip install rouge-score-rs

Rust

cargo add rougers

Change one import

The scorer, the options and the returned Score tuples are the same as in rouge-score:

python
from rouge_score_rs import rouge_scorer  # was: from rouge_score import rouge_scorer

scorer = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL", "rougeLsum"], use_stemmer=True)
scores = scorer.score(
    "The quick brown fox jumps over the lazy dog",
    "The quick brown dog jumps on the log.",
)
scores["rouge1"]
# Score(precision=0.75, recall=0.6666666666666666, fmeasure=0.7058823529411765)

score_multi, custom tokenizers, split_summaries, scoring.BootstrapAggregator and tokenize.tokenize work as before, and scorers can be pickled for worker processes. The migration guide covers dependencies, multiprocessing and how to check parity on your own data.

Score many pairs on all cores

score_batch takes lists of references and predictions, releases the GIL and runs on every CPU core with the default tokenizer:

python
results = scorer.score_batch(references, predictions)  # one dict per pair, in order
results[0]["rougeL"].fmeasure

Set RAYON_NUM_THREADS to limit the threads. With a custom Python tokenizer or split_summaries=True, and in a process forked after a parallel call, pairs are scored one by one.

How much faster

2,000 pairs from each test set, lead sentences as predictions, rouge1, rouge2, rougeL and rougeLsum. Seconds, median of three runs on an Apple M1 Pro. Every precision, recall and F-measure matched rouge-score 0.1.2 exactly.

DatasetStemmingrouge-scorerouge-score-rsscore_batch, 8 threadsSpeedup
CNN/DailyMailno2.5210.0430.011239×
CNN/DailyMailyes5.2720.0930.018300×
XSumno0.3760.0130.00583×
XSumyes1.2190.0250.006196×
PubMedno13.6520.1690.044311×
PubMedyes19.6460.2800.056353×

Through Hugging Face Evaluate, with its default bootstrap aggregation, compute is 3–72× faster, because aggregation and input handling stay in Python. The benchmark page has the method, memory use, a comparison with other ROUGE packages and the commands to reproduce it.

Other languages

The default tokenizer is the one from rouge-score: it keeps only ASCII letters and digits, so Ukrainian, Arabic, Hindi or Chinese text scores zero. Version 0.2 adds two opt-in tokenizers that run natively, also in score_batch:

python
from rouge_score_rs import rouge_scorer, tokenizers

default = rouge_scorer.RougeScorer(["rouge1"])
default.score("Уряд оголосив нові заходи.", "Нові заходи оголосив уряд.")["rouge1"]
# Score(precision=0.0, recall=0.0, fmeasure=0.0)

unicode = rouge_scorer.RougeScorer(["rouge1", "rougeL"], tokenizer=tokenizers.UnicodeTokenizer())
unicode.score("Уряд оголосив нові заходи.", "Нові заходи оголосив уряд.")["rouge1"]
# Score(precision=1.0, recall=1.0, fmeasure=1.0)

chinese = rouge_scorer.RougeScorer(["rouge1", "rougeL"], tokenizer=tokenizers.CharacterTokenizer())
chinese.score("今天天气很好", "今天天气不错")["rouge1"]
# Score(precision=0.6666666666666666, recall=0.6666666666666666, fmeasure=0.6666666666666666)

Hugging Face Evaluate

The SyntaxSpirits/rouge metric takes the same inputs and options as Evaluate's built-in rouge and returns the same scores:

python
import evaluate

rouge = evaluate.load("SyntaxSpirits/rouge")  # was: evaluate.load("rouge")
results = rouge.compute(predictions=predictions, references=references)

It needs pip install evaluate "rouge-score-rs[aggregate]>=0.1.1". Optional backends for the built-in metrics are proposed upstream in huggingface/evaluate#827 and huggingface/lighteval#1426; both are open.

Rust

rust
use rougers::{RougeScorer, RougeType};

let scorer = RougeScorer::new(["rouge1", "rougeL"], true)?; // true = Porter stemming
let scores = scorer.score(
    "The quick brown fox jumps over the lazy dog",
    "The quick brown dog jumps on the log.",
);
let rouge_l = scores.get(RougeType::L).unwrap();
println!("{:.4}", rouge_l.fmeasure); // 0.5882

RougeScorer is Send + Sync, so one scorer can be shared across threads, for example with Rayon.

How parity is tested

The test suite fails on any difference in any float between rouge-score-rs and rouge-score 0.1.2:

Known differences: an invalid ROUGE type raises ValueError when the scorer is created rather than on the first call; texts must be str, not bytes; BootstrapAggregator needs the [aggregate] extra (NumPy) and split_summaries=True the [sentences] extra (NLTK). Without them the package has no Python dependencies.

Questions

Will my reported scores change?

No, with the default tokenizer. The scores are the same floats rouge-score 0.1.2 returns, so numbers in papers and dashboards stay comparable. Only the opt-in Unicode and character tokenizers give different scores.

Do I need a Rust toolchain?

No. PyPI has prebuilt abi3 wheels for Linux (x86_64 and aarch64, glibc and musl), macOS (Intel and Apple Silicon) and Windows x86_64, one wheel per platform for every Python from 3.9. Other platforms build from the source distribution and need Rust.

Is this an official Google project?

No. It is an independent port under the same Apache-2.0 licence as rouge-score and NLTK, whose scoring, tokenization and stemming logic it reproduces.