Quick start
Pass a Matrix or an array of arrays: one row per person, one column per item, 1 for a correct answer, 0 for a wrong one and nil when the item was not answered.
require "irt_ruby"
responses = [
[1, 1, 0, 1],
[0, 1, 0, nil],
[1, 1, 1, 1],
[0, 0, 0, 1]
]
result = IrtRuby::RaschModel.new(responses).fit
result[:abilities] # one value per person, higher = more able
result[:difficulties] # one value per item, higher = harderRasch, 2PL and 3PL
two_pl = IrtRuby::TwoParameterModel.new(responses).fit
two_pl[:discriminations] # how sharply each item separates weaker and stronger people
three_pl = IrtRuby::ThreeParameterModel.new(responses).fit
three_pl[:guessings] # chance of a correct answer by guessing| Model | Item parameters | Use when |
|---|---|---|
RaschModel | difficulty | items should weigh equally, or data are scarce |
TwoParameterModel | difficulty, discrimination (clamped to 0.01–5.0) | some items tell people apart better than others |
ThreeParameterModel | difficulty, discrimination, guessing (clamped to 0–0.35) | multiple-choice items that can be guessed |
Missing answers
Every model takes missing_strategy:
:ignore(default) leavesnilanswers out of the likelihood and gradients;:treat_as_incorrectcounts them as0, for example when a skipped question means the person could not answer it;:treat_as_correctcounts them as1.
model = IrtRuby::TwoParameterModel.new(
responses,
missing_strategy: :treat_as_incorrect,
max_iter: 500,
learning_rate: 0.05,
tolerance: 1e-7,
param_tolerance: 1e-7,
decay_factor: 0.5
)
model.fitHow fitting works
Parameters start at random values and are fitted by gradient ascent on the log-likelihood. If a step lowers the likelihood, it is undone and the learning rate is multiplied by decay_factor. Fitting stops after max_iter iterations or when both the log-likelihood change and the average parameter update fall below tolerance and param_tolerance. Call srand before creating the model if you need the same starting values every run.
Where it helps
- Scoring quizzes and exams in Rails apps by ability rather than by raw percentage.
- Finding items that are too easy, too hard or do not discriminate, before reusing a question bank.
- Analysing which benchmark questions separate stronger and weaker language models, the topic of our study on IRT for budget-constrained LLM evaluation.