Skip to content

Open-source libraries · Rust · Python · Ruby

Tools for evaluation, scientific computing and data applications

We port slow reference implementations to fast, verified code, build gems for AI features in Rails apps, and publish the code and data behind our research.

New · 0.2.1rougers / rouge-score-rs

ROUGE scores identical to rouge-score, 30–80× faster on one core

A Rust port of Google's rouge-score, the implementation behind Hugging Face Evaluate. Every float matches, including the tokenizer, the Porter stemmer and rougeLsum. Change one import and keep your code.

pip install rouge-score-rs
Seconds to score 2,000 PubMed summaries, four ROUGE types with stemming
rouge-score19.65 s
rouge-score-rs, loop0.28 s
score_batch, 8 threads0.056 s

Apple M1 Pro, Python 3.12, median of 3 runs, scores checked bit for bit. Speedups depend on text length, metrics and threads: 30–80× on one core and 83–353× on eight across CNN/DailyMail, XSum and PubMed. Full benchmarks

Libraries

Also maintained here

All 9 libraries →

langdetect-rs

RustPython
New

Language detection identical to Python's langdetect, bit for bit, 40–52× faster on one core and up to 315× with all cores.

PyPI versioncrates.io version
pip install langdetect-rs

rust-lstm

Rust

LSTM, GRU and BiLSTM networks with full backpropagation through time, gradients checked against finite differences and PyTorch.

crates.io versioncrates.io downloadsGitHub stars
cargo add rust-lstm

engram

Ruby
Pre-1.0

Long-term memory for AI agents in Ruby: recall the facts that matter and add them to the prompt, stored in your own Postgres.

Gem versionGem downloads
gem install engram

Research

Code and data you can rerun

Each study ships with its code, prompts, raw outputs and a DOI on Zenodo.

All studies →

Oct 2026 · O. Kholodniak, O. Prokhorov

Gradient reach in recurrent networks: code and training runs

How dropout, zoneout and gate initialisation change how far gradients reach back in time in LSTM and GRU networks, and whether that reach, measured early in training, predicts which runs learn long-range dependencies. Trained with rust-lstm.

Oct 2026 · O. Kholodniak

Comparing IRT models and simple baselines for budget-constrained LLM evaluation

Item response models against simple estimators on public model × item correctness matrices: how well each estimates item difficulty, a new model's capability and its full-benchmark accuracy when only 0.5–80% of the items are run.

Oct 2026 · O. Kholodniak

QoS-aware orchestration of local and cloud large language models in distributed information systems

Choosing per request between a small on-device model, a larger LAN server and a cloud API: latency model, routing policies, a netem testbed and the measurement data.

Updates

Latest releases

Who is behind this

SyntaxSpirits is maintained by Oleksandr Kholodniak, a software engineer and PhD researcher at the National Aerospace University “Kharkiv Aviation Institute” in Ukraine. The libraries come out of production Ruby work and research on evaluating language models.

Found a bug or need a feature? Open an issue on the library's GitHub page, or write to hello@syntaxspirits.dev.