Change one import
Functions, classes, options, exceptions and the returned Language objects are the ones from langdetect 1.0.9:
from langdetect_rs import DetectorFactory, detect, detect_langs # was: from langdetect import ...
DetectorFactory.seed = 0
detect("Ein, zwei, drei, vier")
# 'de'
detect_langs("Война и мир — роман Льва Толстого.")
# [ru:0.5714279654116534, bg:0.42857201882885804]Detector with append, set_alpha, set_prior_map and set_max_text_length, DetectorFactory with load_profile and load_json_profile, LangDetectException with its error codes, and utils.lang_profile.LangProfile work as before. Factories and detectors can be pickled.
Detect many texts on all cores
from langdetect_rs import detect_batch, detect_langs_batch
detect_batch(["Guten Morgen, wie geht es dir heute?", "Bonjour tout le monde", "1234"])
# ['de', 'fr', None]The batch functions release the GIL and use every core. A text without features gives None where detect would raise. RAYON_NUM_THREADS limits the threads; in a process forked after a batch call, texts are detected one by one.
Speed up libraries that import langdetect
When a library imports langdetect itself, call install() before importing it:
import langdetect_rs
langdetect_rs.install() # `import langdetect` now returns langdetect_rs
from unstructured.partition.auto import partitionOn unstructured's example documents this gave the same metadata.languages and made partitioning 2–10× faster for small documents such as e-mails and web pages, where language detection was most of the work. Modules that imported langdetect before the call keep the original.
How much faster
2,000 texts from each test set, milliseconds per text, best of three runs on an Apple M1 Pro with CPython 3.13.
| Data | Seed | langdetect | langdetect-rs | detect_batch, 8 cores | Speedup |
|---|---|---|---|---|---|
| papluca (short texts) | none | 1.452 | 0.0365 | 0.0059 | 40× / 246× |
| WiLI-2018 (paragraphs) | none | 3.455 | 0.0786 | 0.0118 | 44× / 293× |
| papluca (short texts) | 0 | 1.465 | 0.0285 | 0.0047 | 51× / 315× |
| WiLI-2018 (paragraphs) | 0 | 3.446 | 0.0660 | 0.0116 | 52× / 297× |
Starting Python and detecting one text takes 74 ms instead of 244 ms, because there are no JSON profiles to parse. The loaded profiles take about 19 MB instead of 67 MB. Method and raw results: BENCHMARKS.md.
How "identical" is checked
The tests run every scenario on both packages and compare full probability vectors, exceptions and messages, on CPython 3.9 to 3.15. All 127,500 texts of the WiLI-2018 and papluca test sets give the same results on Linux, macOS and Windows. langdetect's own results depend on the interpreter, and langdetect-rs follows it:
- From Python 3.12,
sum()uses compensated summation, solangdetect's probabilities differ between Python 3.11 and 3.12 for 38% of texts, in the last bits. str.isupper()follows the Python's Unicode version; tables for Unicode 13.0 to 17.0 are built in.- Profiles are loaded in
os.listdirorder; withlangdetectinstalled,langdetect-rsuses the same order.
The article Porting langdetect to Rust, bit for bit explains how the random walk is reproduced and where these differences come from.
Known differences
- Without a seed, each detection is seeded from a per-thread generator seeded by the operating system, not from 2.5 KB of
os.urandom. Results are random either way. DetectorFactory.word_lang_prob_mapandDetector.randomdo not exist, andset_verbose()prints nothing.LangDetectExceptionis a separate class fromlangdetect's (afterinstall()they are the same).
Rust
use langdetect::{DetectorFactory, Seed};
assert_eq!(langdetect::detect("Ein, zwei, drei, vier")?, "de");
let mut factory = DetectorFactory::builtin().clone();
factory.seed = Seed::Int(0);
for lang in factory.detect_langs("Ceci est une phrase en français.")? {
println!("{lang}");
}DetectorFactory is Send + Sync. Compat selects which Python's behaviour to reproduce; the default is CPython 3.14.
Apache-2.0, like langdetect
pip install langdetect-rsSource: github.com/SyntaxSpirits/langdetect-rs. An independent project, not affiliated with the authors of langdetect or language-detection.