Skip to content

Language detection identical to langdetect, 40–50× faster

langdetect-rs reproduces Python's langdetect down to the last bit of every probability, random walk included. It is 40–52× faster on one core, 245–315× faster with all eight cores of an M1 Pro, and keeps its profiles in a third of the memory. Change the import and keep the rest of your code.

PyPI version crates.io version CI status

Python 3.9+ · wheels for Linux, macOS and Windows

pip install langdetect-rs

Rust

cargo add langdetect

Change one import

Functions, classes, options, exceptions and the returned Language objects are the ones from langdetect 1.0.9:

python
from langdetect_rs import DetectorFactory, detect, detect_langs  # was: from langdetect import ...

DetectorFactory.seed = 0
detect("Ein, zwei, drei, vier")
# 'de'
detect_langs("Война и мир — роман Льва Толстого.")
# [ru:0.5714279654116534, bg:0.42857201882885804]

Detector with append, set_alpha, set_prior_map and set_max_text_length, DetectorFactory with load_profile and load_json_profile, LangDetectException with its error codes, and utils.lang_profile.LangProfile work as before. Factories and detectors can be pickled.

Detect many texts on all cores

python
from langdetect_rs import detect_batch, detect_langs_batch

detect_batch(["Guten Morgen, wie geht es dir heute?", "Bonjour tout le monde", "1234"])
# ['de', 'fr', None]

The batch functions release the GIL and use every core. A text without features gives None where detect would raise. RAYON_NUM_THREADS limits the threads; in a process forked after a batch call, texts are detected one by one.

Speed up libraries that import langdetect

When a library imports langdetect itself, call install() before importing it:

python
import langdetect_rs

langdetect_rs.install()  # `import langdetect` now returns langdetect_rs

from unstructured.partition.auto import partition

On unstructured's example documents this gave the same metadata.languages and made partitioning 2–10× faster for small documents such as e-mails and web pages, where language detection was most of the work. Modules that imported langdetect before the call keep the original.

How much faster

2,000 texts from each test set, milliseconds per text, best of three runs on an Apple M1 Pro with CPython 3.13.

DataSeedlangdetectlangdetect-rsdetect_batch, 8 coresSpeedup
papluca (short texts)none1.4520.03650.005940× / 246×
WiLI-2018 (paragraphs)none3.4550.07860.011844× / 293×
papluca (short texts)01.4650.02850.004751× / 315×
WiLI-2018 (paragraphs)03.4460.06600.011652× / 297×

Starting Python and detecting one text takes 74 ms instead of 244 ms, because there are no JSON profiles to parse. The loaded profiles take about 19 MB instead of 67 MB. Method and raw results: BENCHMARKS.md.

How "identical" is checked

The tests run every scenario on both packages and compare full probability vectors, exceptions and messages, on CPython 3.9 to 3.15. All 127,500 texts of the WiLI-2018 and papluca test sets give the same results on Linux, macOS and Windows. langdetect's own results depend on the interpreter, and langdetect-rs follows it:

The article Porting langdetect to Rust, bit for bit explains how the random walk is reproduced and where these differences come from.

Known differences

Rust

rust
use langdetect::{DetectorFactory, Seed};

assert_eq!(langdetect::detect("Ein, zwei, drei, vier")?, "de");

let mut factory = DetectorFactory::builtin().clone();
factory.seed = Seed::Int(0);
for lang in factory.detect_langs("Ceci est une phrase en français.")? {
    println!("{lang}");
}

DetectorFactory is Send + Sync. Compat selects which Python's behaviour to reproduce; the default is CPython 3.14.

Apache-2.0, like langdetect

pip install langdetect-rs

Source: github.com/SyntaxSpirits/langdetect-rs. An independent project, not affiliated with the authors of langdetect or language-detection.