Measuring with Text

Text violates conditional independence

  • For CTT and Rasch models good test design can result in items \(\Ym\) not violating assumptions of s.i. conditional on latent variables.

  • For text there does not exist a latent variable explaining the stochastic dependence of words in a text (because text has syntactical structure, among other reasons).

  • We introduce item randomisation to induce stochastical independence required for aggregation.

    Randomisation has a cost: estimation cannot use any attributes of items like difficulty.

Requirements of a Measurement Theory

The introduced theory meets all the requirements of a Measurement Theory.

Text observables, Random experiment

  1. A person \(u \in \Omega_U\) is drawn, with person random variable \(U : \Omega \rightarrow \Omega_U\).

  2. A multivariate observation \(o \in \Omega_O\) is made of this person, consisting of \(m\) words,1

    With manifest random Variables \(Y_i : \Omega \rightarrow \{t_0, t_1,…,t_l\}, i=1, \ldots, m\)

    \(\Omega_O = \Omega_S^m\) (set of possible texts) with \(\Omega_S=\{t_0, t_1,…,t_l\}\) (set of \(l\) possible words in a language).

With the product set \(\Omega = \Omega_U \times \Omega_O \times \Omega_X\), the set of possible outcomes of the random experiment. An observation \(\omega \in \Omega\) is \(\omega=(u,\ym)\).

Text is … complicated

Texts have syntactical structure

  • \(\Ym\) are stochastically dependent.
  • Joint distribution of all possible texts \(\Ym\) is intractable.

→ Assumptions like τ-equivalence are highly unrealistic

Doubts whether computers can understand text

  • “recursive nature of language” (Chompski)

First, stochastic dependence due to syntactical structure needs destroying.


  1. To keep notation here simple, we consider all texts to have length \(m\). In practice texts are of finite length, and one can fill with “empty word” \(t_0\). ↩︎

Gregor Kappler
Gregor Kappler
Independent Researcher and Programmer

My research interests include probability theory, psychometrics, language analysis and programmable ideas.