Measuring with Text
Text violates conditional independence
For CTT and Rasch models good test design can result in items \(\Ym\) not violating assumptions of s.i. conditional on latent variables.
For text there does not exist a latent variable explaining the stochastic dependence of words in a text (because text has syntactical structure, among other reasons).
We introduce item randomisation to induce stochastical independence required for aggregation.
Randomisation has a cost: estimation cannot use any attributes of items like difficulty.
Requirements of a Measurement Theory
The introduced theory meets all the requirements of a Measurement Theory.
Text observables, Random experiment
A person \(u \in \Omega_U\) is drawn, with person random variable \(U : \Omega \rightarrow \Omega_U\).
A multivariate observation \(o \in \Omega_O\) is made of this person, consisting of \(m\) words,1
With manifest random Variables \(Y_i : \Omega \rightarrow \{t_0, t_1,…,t_l\}, i=1, \ldots, m\)
\(\Omega_O = \Omega_S^m\) (set of possible texts) with \(\Omega_S=\{t_0, t_1,…,t_l\}\) (set of \(l\) possible words in a language).
With the product set \(\Omega = \Omega_U \times \Omega_O \times \Omega_X\), the set of possible outcomes of the random experiment. An observation \(\omega \in \Omega\) is \(\omega=(u,\ym)\).
Text is … complicated
Texts have syntactical structure
- \(\Ym\) are stochastically dependent.
- Joint distribution of all possible texts \(\Ym\) is intractable.
→ Assumptions like τ-equivalence are highly unrealistic
Doubts whether computers can understand text
- “recursive nature of language” (Chompski)
First, stochastic dependence due to syntactical structure needs destroying.
To keep notation here simple, we consider all texts to have length \(m\). In practice texts are of finite length, and one can fill with “empty word” \(t_0\). ↩︎