Skip to content

Item response theory in Ruby

irt_ruby fits the Rasch, two-parameter and three-parameter logistic models to right/wrong answers. It estimates how able each person is and how hard, how discriminating and how guessable each item is, from tests, quizzes or benchmark results.

Gem version Gem downloads GitHub stars

Ruby gem · MIT licence

gem install irt_ruby

Quick start

Pass a Matrix or an array of arrays: one row per person, one column per item, 1 for a correct answer, 0 for a wrong one and nil when the item was not answered.

ruby
require "irt_ruby"

responses = [
  [1, 1, 0, 1],
  [0, 1, 0, nil],
  [1, 1, 1, 1],
  [0, 0, 0, 1]
]

result = IrtRuby::RaschModel.new(responses).fit
result[:abilities]     # one value per person, higher = more able
result[:difficulties]  # one value per item, higher = harder

Rasch, 2PL and 3PL

ruby
two_pl = IrtRuby::TwoParameterModel.new(responses).fit
two_pl[:discriminations]  # how sharply each item separates weaker and stronger people

three_pl = IrtRuby::ThreeParameterModel.new(responses).fit
three_pl[:guessings]      # chance of a correct answer by guessing
ModelItem parametersUse when
RaschModeldifficultyitems should weigh equally, or data are scarce
TwoParameterModeldifficulty, discrimination (clamped to 0.01–5.0)some items tell people apart better than others
ThreeParameterModeldifficulty, discrimination, guessing (clamped to 0–0.35)multiple-choice items that can be guessed

Missing answers

Every model takes missing_strategy:

ruby
model = IrtRuby::TwoParameterModel.new(
  responses,
  missing_strategy: :treat_as_incorrect,
  max_iter: 500,
  learning_rate: 0.05,
  tolerance: 1e-7,
  param_tolerance: 1e-7,
  decay_factor: 0.5
)
model.fit

How fitting works

Parameters start at random values and are fitted by gradient ascent on the log-likelihood. If a step lowers the likelihood, it is undone and the learning rate is multiplied by decay_factor. Fitting stops after max_iter iterations or when both the log-likelihood change and the average parameter update fall below tolerance and param_tolerance. Call srand before creating the model if you need the same starting values every run.

Where it helps