Applications

CTT: continuous \(Z\) and \(X\)

Equivalence of Classical Test Theory

With item randomization and definitions of \(\xi\) for the discrete cases, now the application of these definitions to CTT is investigated.

We prove that CTT with \(m\) τ-equivalent (parallel) items is equivalent with the proposed model after a linear transformation.

This investigation helps understanding the similarity with and difference to CTT.

Classical Test Theory with τ-equivalent items

  1. A person \(u \in \Omega_U\) is drawn.

  2. An multivariate observation \(\ym \in \mathbb{R}\) is made of this person \(u\) (manifest random variables \(Y_1,\ldots,Y_m : \Omega \rightarrow \mathbb \mathbb{R}\))

  3. A random variable \(X : \Omega \rightarrow \mathbb \mathbb{R}\) is observed for validation.

    Reflective measurement is a special case when \(X\) is also a τ-equivalent item.

Assumptions \(\Ym : \Omega \rightarrow \mathbb{R}\) are τ-equivalent (parallel)

  1. \(\tau\) -equivalence \(\tau := E(Y_1 | U) = \ldots = E(Y_m | U)\)
  2. Error variables \(\epsilon_i=Y_i-\tau\) are uncorrelated with measurement error \(\sigma_{\epsilon}^2=Var(\epsilon_i)\).
  3. Assume also \(X\) is a τ-equivalent item: \(X=\tau + \epsilon_0\).

Definiton of \(\xi\) and Estimator \(\bar \xi\)

With the newly proposed randomization of manifest variables, let \(Z\) be the randomized measurement as defined in \ref{defZ}.1

Assume the regression \(\xi\) is a linear function2:

\begin{align} \label{cttxialphabeta} \xi = & E(X | Z) = \alpha + \beta \cdot Z \end{align}

We define the \Ym-conditional expectation of \(\xi\)

\begin{align} \bar{\xi} := & E(\xi | \Ym) \\\ \label{cttxibaralphabeta} = & \alpha + \beta \cdot \frac{1}{m}\sum_{i=1}^m Y_{i} \end{align}

The proof of (\ref{cttxibaralphabeta}) is obvious with (\ref{cttxialphabeta}) and marginalisation over \(K\).

Theorem: Equivalence \(\bar{\xi}=\alpha + \beta \bar Y\), with \(\beta=\rho_1\) reliability of one item

The latent variable \(E(\xi|U)\) is related to CTT with \(\bar Y\) (Equation \ref{defYbar}):

  1. estimates of person scores \(\bar \xi\) are linearly transformed \(\bar Y\) with a slope identical to reliability \(\beta=\rho_1\):

    \begin{align} \bar{\xi} = & \alpha + \beta \bar{Y} \\\ \beta = & \frac{Cov(X,Z)}{Var(Z)}\\\ = & \rho_{1} \\\ \alpha = & E(X) - \beta E(Z) \\\ = & \mu \cdot (1-\beta) \end{align}

  2. identical reliability \(\rho_{\bar{Y}}=\rho_{\bar{\xi}}\)

Randomized Variable \(Z\)

\begin{align} \label{cttEZ} E(Z) = & \mu \\\ \label{cttVarZ} Var(Z) = & Var(\tau) + \sigma_{\epsilon}^2 \\\ \label{cttCovXZ} Cov(X,Z) = & Var(\tau) \end{align}

Proofs obvious.

Regression random variable \(\xi\)

\begin{align} E(\xi) = & \alpha + \beta \cdot \mu \end{align}

Proof obvious.

\begin{align} Var(\xi) = & \beta^2 Var(Z) \end{align}

Proof:

\begin{align} Var(\xi) = & Var(\alpha + \beta \cdot Z) \\\ = & E\left( (\alpha + \beta \cdot Z -E(\alpha + \beta \cdot Z))^2 \right) \\\ = & \beta^2 E\left( (Z -E(Z))^2 \right) \\\ = & \beta^2 Var(Z) \end{align}

\begin{align} E(\xi | U) = & \alpha + \beta \cdot \tau \end{align}

Proof:

\begin{align} E(\xi | U) = & E \left[ E(\alpha + \beta \cdot Z | U) \right] \\\ = & \sum_{i=1}^m P(K=i | U) E(\alpha + \beta \cdot Z | K=i, U) \\\ = & \sum_{i=1}^m \frac{1}{m} \left( \alpha + \beta \cdot \underbrace{E(Y_i | U)}_{=\tau} \right) \\\ = & \alpha + \beta \cdot \tau \end{align}

With \(K\) being a uniform random variable for the randomization of items (Equation \ref{defK}).

Estimator \(\bar \xi\)

\begin{align} E(\bar{\xi}) = & E(\alpha + \beta \cdot \frac{1}{m}\sum_{i=1}^m Y_{i})\\\ = & \alpha + \beta \mu.\\\ E(\bar{\xi}|U) = & E(\alpha + \beta \cdot \frac{1}{m}\sum_{i=1}^m Y_{i} | U)\\\ = & \alpha + \beta \tau = E(\xi | U).\\\ \label{cttVarxibar} Var(\bar{\xi}) = & \beta^2 Var(\tau) + \frac{\beta^2}{m} \sigma_{\epsilon}^2\\\ \label{cttVarxibarGU} Var(\bar{\xi} | U) = & \frac{\beta^2}{m} \sigma_{\epsilon}^2 \end{align}

Proofs Estimator \(\bar \xi\)

Proof of (\ref{cttVarxibar}):

\begin{align} Var(\bar{\xi}) = & Var \left( \alpha + \beta \cdot \frac{1}{m}\sum_{i=1}^m Y_{i} \right) \\\ = & \beta^2 Var \left( \sum_{i=1}^m \frac{1}{m}(\tau + \epsilon_i) \right) \\\ = & \beta^2 Var\left(\tau + \frac{1}{m}\sum_{i=1}^m \epsilon_i \right) \\\ = & \beta^2 Var(\tau) + \frac{\beta^2}{m} \sigma_{\epsilon}^2 \end{align}

The last identitly follows from the fact that \(\tau, \epsilon_1, \ldots, \epsilon_m\) are stochastically independent and \(Var(\epsilon_i)=\sigma_{\epsilon}^2\).

Proof of (\ref{cttVarxibarGU}):

\begin{align} Var(\bar{\xi} | U) = & E \left[ \left( \bar{\xi} - E(\bar{\xi} | U) \right)^2 \right] \\\ = & E \left[ \left( \alpha + \beta \frac{1}{m} \sum_{i=1}^m Y_i - \alpha - \beta \tau \right)^2 \right] \\\ = & \beta^2 E \left[ \left( \frac{1}{m} \sum_{i=1}^m Y_i - \tau \right)^2 \right] \\\ = & \beta^2 E \left[ \left( \frac{1}{m} \sum_{i=1}^m (\tau + \epsilon_i) - \tau \right)^2 \right] \\\ = & \beta^2 E \left[ \left( \frac{1}{m} \sum_{i=1}^m \epsilon_i \right)^2 \right] \\\ = & \frac{\beta^2}{m^2} Var \left[ \sum_{i=1}^m \epsilon_i \right] = \frac{\beta^2}{m} \sigma_{\epsilon}^2 \end{align}

The last identitly follows from the fact that \(\epsilon_1, \ldots, \epsilon_m\) are stochastically independent and \(Var(\epsilon_i)=\sigma_{\epsilon}^2\).

Reliability

\(\bar{\xi}\) and \(\bar Y\) have equal reliability.

Proof:

\begin{align} \rho_{\bar{\xi}} = & 1- \frac{Var(\bar{\xi} | U)}{Var(\bar{\xi})} \\\ = & 1-\frac{\frac{\beta^2}{m} \sigma_{\epsilon}^2}{\beta^2 Var(\tau) + \frac{\beta^2}{m} \sigma_{\epsilon}^2}\\\ = & 1-\frac{\frac{1}{m} \sigma_{\epsilon}^2}{Var(\tau) + \frac{1}{m} \sigma_{\epsilon}^2}\\\ = & \frac{Var(\tau)}{Var(\tau)+\frac{i}{m}\sigma_{\epsilon}^2} \\\ = & \rho_{\bar{Y}} \end{align}

Discrete \(Z\), continuous \(X\)

Estimating latent variables \(\bar \xi\) from text

In the case of finite possible outcomes for \(Y_i\), given an observation \(\omega\), \(\bar{\xi}(\omega)\) is an unbiased estimate of \(E(\xi|U=U(\omega))\), see Equation (\ref{defxibar}).

\begin{align} \bar{\xi} = & \sum_{z \in \Omega_S} E(X | Z=z) \cdot J_z \end{align}

FRONT \(Z=z\) -conditional expectation of \(X\)

\begin{align} E(X | Z=z) & = \frac{ E \left[ X \cdot J_{Z=z} \right]}{ E(J_{Z=z}) } \end{align}

Proof by marginalisation over all \(Y_i\) on next slide!

VERBESSERN Proof: \(Z=z\) -conditional expectations \(X\)

By construction \(Z\) is \(\mathbf{Y}\) -conditionally independent from \(X\):

\begin{align} \label{zycondiidx} P(Z=z | X=x, \mathbf{Y=Y}(\omega)) &= P(Z=z | \mathbf{Y=Y}(\omega)) \\\ \label{defPz} &= \frac{1}{m} \sum_{i=1}^m I_{Y_i=z}(\omega) \\\ &= \frac{1}{m} \sum_{i=1}^m \delta(Y_i(\omega),z) \end{align}

With vector notation \(\mathbf{Y} := (\Ym)\) and \(\mathbf{Y}(\omega)=(\ym)\), and with Kronecker’s \(\delta(a,b)=1\) if \(a=b\) and \(\delta(a,b)=0\) if \(a \ne b\).

The joint probability of \(X\) and \(Z\) is by the law of total probability, over all possible events in \(\mathbf{y} \in \Omega_O=\mathbf{Y}(\Omega)\):

\begin{align} P(X=x, Z=z) & = \sum_{\mathbf{y} \in \Omega_O} P(Z=z, X=x, \mathbf{Y=y}) \\\ & = \sum_{\mathbf{y} \in \Omega_O} P(Z=z | X=x, \mathbf{Y=y}) \cdot P(X=x, \mathbf{Y=y}) \\\ & \stackrel{(\ref{zycondiidx})}{=} \sum_{\mathbf{y} \in \Omega_O} P(Z=z | \mathbf{Y=y}) \cdot P(X=x, \mathbf{Y=y}) \\\ \label{jointPXZ} & \stackrel{(\ref{defPz})}{=} \sum_{\mathbf{y} \in \Omega_O} \left( \frac{1}{m} \sum_{i=1}^m \delta(y_i,z) \right) \cdot P(X=x, \mathbf{Y=y}) \end{align}

\begin{align} E(X | Z=z) & = \frac{\int_{-\infty}^{\infty} x \cdot P(X=x,Z=z) dx}{P(Z=z)} \\\ & \stackrel{(\ref{jointPXZ})}{=} \frac{1}{P(Z=z)} \int_{-\infty}^{\infty} x \cdot \left[ \sum_{\mathbf{y} \in \Omega_O} \left( \frac{1}{m} \sum_{i=1}^m \delta(y_i,z) \right) \cdot P(X=x, \mathbf{Y=y}) \right] dx \\\ & = \frac{1}{P(Z=z)} \left[ \sum_{\mathbf{y} \in \Omega_O} \left( \frac{1}{m} \sum_{i=1}^m \delta(y_i,z) \right) \cdot \int_{-\infty}^{\infty} x \cdot P(X=x, \mathbf{Y=y}) dx \right] \\\ & = \frac{1}{P(Z=z)} \sum_{\mathbf{y} \in \Omega_O} \left( \frac{1}{m} \sum_{i=1}^m \delta(y_i,z) \right) \cdot E(X | \mathbf{Y=y}) P(\mathbf{Y=y}) \end{align}

Conditional on \(\mathbf{Y}=\mathbf{y}=(y_1, \ldots, y_m)\) the Kronecker’s \(\delta(y_i,z)\) are constant and can be pulled into the expectation:

\begin{align} E(X | Z=z) & = \frac{1}{P(Z=z)} \sum_{\mathbf{y} \in \Omega_O} \left( \frac{1}{m}\sum_{i=1}^m \delta(y_i,z) \right) \cdot E(X | \mathbf{Y=y}) \cdot P(\mathbf{Y=y}) \\\ & = \frac{1}{P(Z=z)} \sum_{\mathbf{y} \in \Omega_O} E \left[ \frac{1}{m}\sum_{i=1}^m \delta(y_i,z) \cdot X | \mathbf{Y=y} \right] \cdot P(\mathbf{Y=y}) \end{align}

\normalsize Also the substitution \(\delta(y_i,z)=I_{Y_i=z}\) is correct in the conditional expectation:

\begin{align} E(X | Z=z) & = \frac{1}{P(Z=z)} \sum_{\mathbf{y} \in \Omega_O} E \left[ \frac{1}{m}\sum_{i=1}^m I_{Y_i=z} \cdot X | \mathbf{Y=y} \right] \cdot P(\mathbf{Y=y}) \\\ & \stackrel{(\ref{defJ})}{=} \frac{1}{P(Z=z)} \sum_{\mathbf{y} \in \Omega_O} E \left[ J_{Z=z} \cdot X | \mathbf{Y=y} \right] \cdot P(\mathbf{Y=y}) \\\ & = \frac{E[ J_{Z=z} \cdot X ]}{P(Z=z)} \qed \end{align}

The last equation follows from law of total expectation.

Discrete \(Z\) and \(X\)

VERBESSERN Person-Conditional Regression of \(X\) on \(Z\):

In the case of discrete \(X\), the regressions of \(I_{=xX}\) on a randomly selected item from the observation \(Z\) are random variables, again.

\begin{align} \label{defxi} \xi_x := & E(I_{X=x} | Z) \end{align}

These regressions reflect what we can learn about the validity criterion \(X\) from a single outcome of the test – if we do not know what item the outcome was for.

The latent variable of interest is the \(U\) -conditional expectation of this regression of \(X\) on \(Z\)

\begin{align} \label{defxigu} \label{defvee} E(\xi_x | U) = & E [ E(I_{X=x} | Z) | U ]\\\ = & \sum_{z \in \Omega_S} P(X=x | Z=z) \cdot \zeta_z \end{align}

The proofs in the continuous case apply, with \(P(X=x | Z=z)=E(I_{X=x} | Z=z)\).

IRT

FRONT Simulation: model estimates display nearly identical correlations

Simulation for the Rasch model of IRT and \(X\) being another item with difficulty 0 (a reflective variable) have been done.

The logits of estimates of latent variables \(\bar\xi\) correlate as well with the true person abilities as Rasch estimates (R package eRm).

Table 1: Simulation Parameters:
N1000
person abilities\(\tau \sim N(0, 2)\)
item difficulties\(\theta=(-1,-0.5,0,0.5,1)\)
Table 2: Correlations:
tauymeanxibarrasch
tau1.00000.83380.83470.8355
ymean0.83381.00000.99990.9995
xibar0.83470.99991.00000.9998
rasch0.83550.99950.99981.0000

Simulation: model estimates rescaled (more items)


  1. Selection with independent randomization variable \(K : \Omega \rightarrow \{1, \ldots, m \}\) uniform. ↩︎

  2. This is true if \(X\) is a τ-equivalent item, in the general case I think this assumptions is not required… ↩︎

Gregor Kappler
Gregor Kappler
Independent Researcher and Programmer

My research interests include probability theory, psychometrics, language analysis and programmable ideas.