Porting langdetect to Rust, bit for bit
langdetect has not had a release since 2021 and is still downloaded about ten million times a month. I ported it to Rust so that it returns the same probabilities to the last bit, random walk included. Along the way it turned out that the original does not even agree with itself across Python versions.
11 October 2026 · Oleksandr Kholodniak
Why "the same" had to mean the same bits
langdetect is a Python port of Nakatani Shuyo's Java language-detection. Document pipelines use it to tag text with a language; unstructured, for example, depends on it for metadata.languages. The IFEval checks in lm-eval, lighteval and inspect-evals call it to check the language of model answers. It is pure Python, so detection costs 1–3 ms per text, and loading its 55 language profiles takes a noticeable part of a second.
Faster detectors exist, but they give different answers. If a pipeline filters a corpus, tags documents or scores a benchmark with langdetect, switching to another detector changes its output, and nobody can tell in advance by how much. A replacement is only safe if every result stays the same, and that includes langdetect's probabilities, which come from a random process.
What langdetect actually computes
Each language profile holds frequencies of character 1-, 2- and 3-grams. For a text, langdetect normalises the characters, extracts the n-grams that occur in some profile and then runs seven trials of a random walk:
- Start with equal probabilities for all 55 languages and draw a smoothing value from a normal distribution.
- Pick a random n-gram from the text and multiply each language's probability by how likely that n-gram is in the language.
- Every five steps, normalise the probabilities; stop when one language exceeds 0.99999, or after 1,000 steps.
The result is the average over the trials. Both the n-gram choices and the smoothing come from Python's random module, seeded for each detection: with DetectorFactory.seed = 0 the result is reproducible; without a seed it is drawn fresh from the operating system.
Reproducing CPython's random module
To get the same walk, the Rust code has to draw the same numbers. That means CPython's Mersenne Twister with its init_by_array seeding, random() built from two 32-bit outputs, and the exact paths langdetect takes through the module:
random.choicegoes through_randbelow, which drawsgetrandbits(k)for the bit length of n and rejects values that are too large. Using a modulo instead would pick different n-grams.random.gaussuses the Box–Muller transform and keeps the second value for the next call. Each trial draws one value, so trials alternate between a freshly computed cosine and the cached sine. Seeding clears the cache.- Seeds can be integers, strings, bytes or floats, and CPython turns each type into a state differently (strings are hashed with SHA-512, floats with
hash()). Rather than reimplement that, the Python package lets CPython seed arandom.Randomonce and passes the resulting 624-word state to Rust.
The floating-point arithmetic has to follow the original too. langdetect computes prob[i] *= weight + p[i] for every language, including those where p[i] is zero, and normalises by dividing by the sum. Rust performs exactly these operations in this order: no fused multiply-add, no multiplying by a reciprocal instead of dividing. The cosine, sine and logarithm in gauss come from the platform's maths library on both sides, so CI checks the results on Linux, macOS and Windows, and with musl.
The interpreter is part of the algorithm
Three things outside langdetect's own code decide its output.
sum() changed in Python 3.12
Python 3.12 switched the built-in sum() for floats to compensated (Neumaier) summation. langdetect normalises with sum(prob), so the same text, with the same seed, gets different probabilities on Python 3.11 and 3.12:
from langdetect import DetectorFactory, detect_langs
DetectorFactory.seed = 0
detect_langs("Ceci est une phrase en français.")
# Python 3.11: [fr:0.9999956302010568]
# Python 3.13: [fr:0.9999956302010566]On the 127,500 texts of the WiLI-2018 and papluca test sets, the probabilities differ between the two summation modes for 38% of texts. The differences are in the last bits, and in this sample the detected language never changed, but anyone who stores probabilities or compares runs across environments should know that the Python version is an input. langdetect-rs picks the summation mode from the running interpreter, so it matches langdetect on whichever Python it runs.
str.isupper() follows the Unicode version
To skip acronyms, langdetect ignores n-grams inside runs of capital letters, using str.isupper(). Which characters are uppercase depends on the Unicode version bundled with Python: 13.0 in Python 3.9 and 3.10, up to 17.0 in Python 3.15. Rust's char::is_uppercase follows the Unicode version of the compiler and differs from CPython on 28 to 95 code points, depending on the Python. langdetect-rs ships tables generated from each CPython from 3.9 to 3.15 instead.
Profile order comes from os.listdir
langdetect loads its profiles in os.listdir order, which is not sorted and depends on the file system. The order changes the order of the additions in sum(), and with it the last bit of a probability in about one text in 100,000. When langdetect is installed next to it, langdetect-rs uses the same directory order; otherwise it uses alphabetical order.
One more detail is reproduced on purpose. When langdetect decides whether a text is mostly non-Latin, it compares a Unicode block number with a block name, a comparison that is never equal. As a result, every character from U+0300 up counts as non-Latin. Fixing it would change results, so the port keeps it.
Without a seed, results change between runs
Without a seed, each detection uses a fresh random seed, and the detected language can change from one call to the next. On the same 127,500 texts, two unseeded runs disagreed on the top language for 5.7% of them, mostly very short texts and languages langdetect has no profile for. For the IFEval checks I looked at the 541 official GPT-4 responses: the 95 language checks gave the same verdict in 200 out of 200 runs each, so there it does not matter. If you filter or tag data with langdetect and need the same result twice, set DetectorFactory.seed.
Making it fast
The first Rust version was about 25 times faster than Python. Profiling showed that a third of the time went into preparing the walk rather than the walk itself, and that each step cost about 100 ns, several times what the arithmetic needs. My first guess, that mixing scalar writes with vector reads of the same array stalled the CPU, did not hold up: rewriting for that gave little. What helped:
- a dense row of 55 probabilities for each n-gram, built the first time the walk draws it, so each step is one vectorised loop over the languages;
- a direct lookup table for single characters and reused per-thread buffers instead of hash maps and allocations per call;
- for unseeded detection, a per-thread generator seeded once from the operating system (and again after
fork), instead of reading 2.5 KB of system entropy for every text, as CPython'srandom.seed(None)does.
On an M1 Pro, with 2,000 texts from each test set, langdetect-rs is 40–52 times faster on one core and 245–315 times faster with detect_batch on eight cores. Starting Python and detecting one text takes 74 ms instead of 244 ms, and the loaded profiles take about 19 MB instead of 67 MB. The benchmark page has the method and raw numbers.
Checking that nothing changed
The test suite runs every scenario on both packages and compares the full probability vectors, the exceptions and their messages: sentences in many languages, synthetic texts for all 55 profiles, random Unicode including lone surrogates, property-based tests and every detector option, on CPython 3.9 to 3.15. A separate script compares all 127,500 texts of WiLI-2018 and papluca; it found no differences on Linux, macOS and Windows with Python 3.11 and 3.13, and locally on 3.9 and 3.14.
"Drop-in" also covers the odd corners, because code in the wild reaches them. In langdetect, loading profiles into a factory that already has some writes the new probabilities into the old languages' slots; a prior map of the wrong length raises IndexError after part of the work; a failed detection leaves a partly filled langprob that the next call returns. langdetect-rs does the same in each case.
Using it
from langdetect_rs import DetectorFactory, detect, detect_langs, detect_batch # was: from langdetect import ...
DetectorFactory.seed = 0
detect("Ein, zwei, drei, vier") # 'de'
detect_batch(["Guten Morgen", "Bonjour", "1234"]) # all cores, None where langdetect raisesWhen a library imports langdetect itself, langdetect_rs.install() makes import langdetect return the Rust version. With it, unstructured's partitioners produced the same metadata.languages on its example documents and ran 2–10 times faster on small documents, where language detection was most of the work. I have proposed using it there when installed in Unstructured-IO/unstructured#4547.
The Rust crate is langdetect on crates.io, with the same detector, profiles and a Compat setting that chooses which Python's behaviour to reproduce.
Try it on your data
pip install langdetect-rsGuide and API: langdetect-rs. Source: github.com/SyntaxSpirits/langdetect-rs. If any text gives a different result from langdetect, that is a bug: please open an issue.