What is measuring?
Several observations on a unit are combined to estimate an attribute as number(s) with meaning in reality.
What is measuring?
Elements of a Measurement/Assessment
- Observations are assessed in order to
- measure some attribute (e.g. mass, intelligence)
- of a unit (a person or an object).
Typically
- measurements are slightly off the “true” value of the attribute: measurement error.
- Attributes change in the course of time.
Examples
Hubble Deep Field (1995)
- 10 day sensor-exposure to an unknown and dark part of the universe1.
- 3000 new galaxies were discovered.
- Some attributes of these galaxies were estimated from the observed images.
Intelligence Testing
- Item structure: Rasch/CTT model assumptions
- Items are thoroughly designed and refined
- Controlled assessment procedure
Measurement and the Social Sciences
“Hard Sciences”
Direct correspondence of assessed data and attribute.
Numerical attributes in physics, chemistry, biology like length, weight, concentration.
“Straight forward” assessment procedures exist
rather clear ontology:
length exists independent of assessment.
“Soft Sciences”
Assessed data are rather “hints” to an attribute.
Observed data is complicated (e.g. behavior, texts)
Designed tests have unclear ontology
“Intelligence is what the intelligence test measures.” – but if intelligence exists outside of the assessment, how is it related to the measurement.
Measurement in social sciences is a challenge.
FRONT History of Measurement in the Social Sciences
Current Theories
- Lord & Novick: Classical Test Theory
- Rasch: Item Response Theory
- Steyer: Latent state-trait Theory
Theories rely on observing items of tests with assumptions of conditional independence.
START Reflective Measurements
- describe in terms of probability theory the expected values of observations for a unit,
- ideally assume that several observations are due to the same one (or few) latent dimensions
- provide means to test assumptions and aggregate information from observations
- need external validation.
Measurement Theory: Random Experiment
Assessment can be described mathematically in terms of probability theory as a random experiment (Steyer):
A person \(u \in \Omega_U\) is drawn from a population (manifest unit random variable \(U : \Omega \rightarrow \Omega_U\)).
An observation \(y \in \Omega_o\) is made of this person \(u\) (manifest random variables \(Y : \Omega \rightarrow \Omega_S\)) 2.
Examples:
- CTT: real-valued \(\Omega_S=\mathbb{R}\)
- IRT, Rasch: \(\Omega_S=\{0, 1\}\)
- IRT, partial credit: \(\Omega_S=\{1,\ldots,k\}\)
Given a probability space \((\Omega, \mathcal{A}, P)\).
Measurement vs. Compression
Not every procedure to compute a number from observations for a unit is a good measurement:
| Latent Variable | Principal Component Analysis | |
|---|---|---|
| example | IQ Test | machine learning |
| utility | theoretical | practical |
| interpretation | measurement | compression |
Requirements of a Measurement Theory
A measurement theory must provide
- random experiment
- a formal description of the assessment process in terms of probability theory,
- well-defined latent variable
- in terms of a person-conditional expectation of a random variable3,
- unbiased estimators for latent variables
- from several assessed observations \(\Ym\) for a unit \(u\), and
Otherwise precise interpretations are not possible.
From well defined latent variables one can derive statistical means to test substantive hypotheses on latent variables which
- abstract from the observations and
- attenuate measurement error.
A Measurement Theory can make Assumptions
A measurement theory can assume a statistical model for several observations \(\Ym\) conditional on the latent variable.
In modern theories this is the case:
- CTT
- τ-equivalent, essential τ-equivalent or τ-congeneric models
- IRT Rasch model
- uni-dimensionality, local stochastic independence.
If a measurement theory makes assumtions, then also statistical tests for model assumptions must be provided.4
Text data in “soft sciences”
- very complicated syntactical structure,
- violates assumptions of conditional independence.
There are doubts whether text-based measurement is possible.
Psychometrics with observed text
Text data in “soft sciences” is still considered useful:
- Qualitative research: hermeneutic text interpretation
- Expert ratings: subsequent rating-based statistics
- Statistical/computational text models.
Goal: Psychometrics with Text
We present a well-defined measurement theory of text assessments:
- well-defined in terms of a) random experiment, b) latent variables and c) unbiased estimators
- but is failing the subtle meanings in a text.5
Skript
Eigentlich gilt mein Interesse formalen Sprachen und der Literatur. Gemäß des Beschreibungsinhalts nahezu aller Literatur interessiert mich besonders der Mensch, das dahinterliegende Objekt, und daher Psychologie im weitesten Sinne.
In meinem kürzest zurückliegenden Versuch, die Analyse von Texten in der Psychologie zu etablieren, begab ich mich an einen Lehrstuhl für Psychometrie.
Als Psychometriker sind wir bemüht, Aussagen über den Menschen aufgrund von Beobachtungen machen zu können – spezifisch: mathematisch zu gewährleisten, dass die latenten Variablen, mit welchen wir Personen beschreiben, präzise mathematische Bedeutung haben, und in den Beobachtungen verankert sind. Nur dann können alle Untersuchungen auf Basis dieser Variablen eindeutig interpretiert werden. Natürlich ist das ein sehr hoher Anspruch, und
in Wahrheit werden viele wissenschaftlichen Studien innerhalb der Psychologie diesem Anspruch tatsächlich nicht gerecht.
Es gibt, wie in allen Disziplinen, pragmatische Schulen einerseits und rigorose Schulen andererseits. Die Tradition der Psychometrie baut auf Strukturgleichungsmodellen oder grafischen Modellen auf, und die Konstruktion von Messinstrumenten, die psychometrisch valide sind, gilt im allgemeinen als große Kunst, bei der viel Sorgfalt auf die Auswahl und Formulierung der Fragen in einem Fragebogen zu legen ist.
Es gibt Probleme: Beispiel Desirees Masterarbeit.
Rasch-Modelle sind vor dem Hintergrund der gewaltigen Komplexität des Menschen eine große Herausforderung.
So erscheint ein Unterfangen, Psychometrie mit freiem Text als manifester Variable zu unternehmen, vielleicht erst einmal irrsinnig.
Ich bin also an die mir als am rigorosesten bekannte Institution gegangen, in der Absicht, mich meinerseits zu bemühen, ob auf Basis von Textbeobachtungen solch klar definierte Variablen einer Person abgeleitet werden können, oder irgendeiner anderen beobachteten Einheit wie einen Artikel Review oder ähnlichem.
Was soll ich sagen, es war eine harte Schule! Es ist eine harte Schule. Es herrscht allgemein die Meinung, mit Text könne man vielleicht allgemein ein bisschen herum rechnen, aber latente Variablen zur Personenbeschreibung im rigorosen Sinne seien nicht konstruierbar.
Doch genug von mir, ich nahm viele Anläufe, und Text ist in der Tat von einer Komplexität, die ich nicht nur einmal unterschätzt habe in meinen Ansätzen und Bemühungen. Es zeigte sich, dass in der Anwendung, wie beispielsweise bei den Vorhersagen meine Notizen Überschriften, Probleme zeigten bei Texten mit sehr vielen Texten oder mit sehr wenigen Worten. Es gab heuristische Lösungen, mit fachbegrifflichen Namen wie Prior Verteilungen, die ihrerseits zwar halfen aber in halfen sie für lange Texte waren sie unbrauchbar für kurze Texte, halfen sie für kurze Texte waren sie unbrauchbar für lange Texte.
Häufig Dachte ich, Heureka, ich hab’s. Häufig dachte ich dies in den letzten 20 Jahren. Für mein Interesse an Sprachen las ich das Rätsel der Mythos von Sisyphos von Camus, und jedes Heureka ist eben ein Erreichen des Gipfels gewesen.
Ich experimentierte und promovierte und holte mir Meinungen, und es fanden sich immer wieder Fehler. Meist war es nicht angezeigt, sich im Beweisen zu verstricken, die aufgrund der absehbaren Probleme bei der Programmierung aussichtslos waren, denn beim beweisen kann man sich auch verbeißen.
Nun ist wieder eine Zeit des Heureka.
Also begrüße ich erneut die Zeit des Scheiterns und weiteren Bemühens.
Insofern und wie es wohl bei Sisyphos ist, ist klar, dass dies nicht das Ende der Reise sein kann, und dass, wiewohl ich vielleicht ein Ziel erreicht habe, das überhaupt zu verfolgen anderen vollkommen irrsinnig erschien, für mich ist es von vornherein ein Scheitern an meinem Anspruch, Sprache und Texten tatsächlich gerecht zu werden.
Doch nun zum gemachten!
Current statistical/computational text analysis
Latent Semantic Indexing (LSI)
- singular value decomposition for texts,
- extracted dimensions have no meaning (no latent variables)
Topic Modeling/Deep Learning
- “Latent” variables explaining the data (compression)
- without clear interpretation
Approaches provide no measurement theory for texts!
- Reading the tea leafs criticism applies to all.
- predictions are no person-conditional expectations.
No measurement theory
“Latent” variables that are not well defined
Sometimes “latent” variables are defined by models compressing data (e.g. exploratory factor analysis), not as person (\(U=u\))-conditional expectations of a random variable.
These “latent” variables do not assure well-defined interpretation: what attribute of \(u\) do they reflect (Reading the tea leafs criticism, LDA)?
Violated Model Assumtions
Assumptions of a measurement theory may be violated.
Since typically we test whether model assumption null-hypotheses must be rejected based on the observations, these tests more likely as sample size increases (and as the number of observed variables increases).
START No measurement theory
More on requirements of latent variables?
- person-conditional expectation of a random variable
- that random variable must first be uniquely defined in terms of observations
Models with latent variables explaining the data are most of the time not identified.
- Exploratory Factor analysis: rotations
VERBESSERN From Classical Test Theory to a Text Theory
Outline:
Review of Classical Test Theory introduces concepts (for simplicity with one observed variable).
Then concepts are extended to a new theory
- Removing stochastical structure with randomisation after the observation.
- Introducing a well-defined (formative) model for measurement.
These ideas apply for all kinds of observations, real-valued (CTT), dichotmous (IRT) as well as text.
Real-valued observations: Equivalence of the new Theory with CTT (for simplicity assumtions of τ-equivalence).
IRT and Text observations
Proofs for the diligent proofs are in the appendix.
which many argued was empty. ↩︎
Generally, a test of length \(m\) is assessed, consisting of \(m\) items \(y_1, \ldots, y_m \in \Omega_S\), see below. ↩︎
or a function of such an expectation ↩︎
These tests typically are of the form of null-hypotheses that the assumptions are not violated. Generally it is accepted to suffice if tests fail to reject these null-hypotheses. ↩︎
(sometimes you need to break something to make something.) ↩︎