PowerPoint-Präsentation
Download
Report
Transcript PowerPoint-Präsentation
ITC Conference, Winchester, 2002
Computer-based Testing
Usability of Psychometric Admeasurements
Dr. J. M. Müller
University of Tübingen, Germany
http://www.joergmmueller.de/default.htm
Overview
1.
Introduction: Formal test descriptions in practice
2.
Definition of usability in the context of test
description
3.
Illustrating problems: Reliability
4.
Criteria of usability: foundation, scaling, general
attributes
5.
Two examples of enhanced usability: NDR and PDR
6.
Summary
Introduction: Psychometric admeasurements
in practice today and tomorrow
1.
Test users often use poor quality tests (e.g. Piotrowski et al.; Wade & Baker,
1977) Psychometric knowledge (Moreland et al. 1995)/Competence approach (Bartram, 1995,
1996)
2.
What should be described? CBT: Criteria for software usability
(ISO 9241/10, 1991; Willumeit, Gediga & Hamborg, 1995) and further criteria: platformindependence, possibility of making own norm banking, protection)
3.
How should it be described?
4.
“Good practice” guidelines and standards are based on quality criteria
(e.g. Standards for educational and psychological Testing, PA, 1999; International Guidelines
for testing, ITC, 2000)
Quality Supply
Quality Demand
Definition of Usability
Scope of usability: Usability in the context of psychological testing concerns all
important kinds of information for test users to describe a test for various
purposes and the ways to communicate them. This includes test manuals as
well as a formal test descriptions with the help of psychometric
admeasurements.
Aim of usability: The product or effect of good usability is that any test user finds all
necessary information quickly and in a proper standardized form, ready to use
for answering the questions of the test users to enable them to decide whether
a test is an appropriate help for the diagnostic question.
Frame of usability: Quality assurance in the context of psychological testing refers
to test construction, test translation, test description and the use of tests in
practice.
Methods to enhance quality control can contain guidelines for test use,
standards for test description, etc. Usability is a strategy to enhance quality on
the level of formal description.
Consequences of usability concern the reengineering of formal test description,
Indices of measurement of error
CTT
Generalizability Theory
Dimensional construct
IRT
nonspecific
misclassification
specific
misclassification
Categorical construct
Measurement of error
Relationships between indices
of error of measurement
Y/ Kappa/ Phi
Reliability
Phi
Standard error score
Korrelation
Kappa
Top-down vs. bottom-up strategy
to develop a coefficient
Practitioner‘s point of view
Interpretation of the score
(operational meaning)
Defining the operational
meaning
scale (correction)
Scale definition
Algorithm
Index: Defining the
influencing factors
Index: Generic formula
Index: Generic formula
test theory/statistic
Specification of within
a test theory
Scientist‘s point of
Rescaling reliability:
Number of distinctive results (NDR)
(Wright & Master, 1982; Lehrl & Kinzel, 1973; Müller, 2001)
k x2 x1 0.05 1,96 sx 21 rtt
Formula
Test score distribution
Rang R
x1
x2
critical
critical
critical
critical
critical
difference difference difference difference difference
R
2
D
k 2 * 1 rtt 12
R = test score range
k = critical
difference
Criteria of usability
for formal quality criteria
(modified from Müller, 2001, 2002a,b; Goodmann & Kruskal, 1954)
Foundation
1.
2.
3.
4.
5.
Unambiguous
operational
meaning
Unambiguous
formal definition
Broad application
area
Relevant
dependencies
Independent of
irrelevant factors
Scale Definition
1.
2.
3.
Meaningful scale unit,
that implies:
• Interval scale
• Positive values
• Defined range of
values
Comparable to the
reference scale
Significant scale unit that
implies a minimum of
observations (Nmin)
Global attributes in
using
1.
2.
3.
4.
5.
6.
Relevance
Informative (not
redundant)
Predictable for the test
user (nominal/actual
value comparison)
Easy to learn
Easy to utilise
Fisher(1925) criteria
of estimating
NDR at work...
NDR = 2
NDR = 5
NDR = 10
r = .50
r = .92
r = .98
Distribution of reliability coefficient
Distribution of NDR coefficient
Conclusion: many precise tests
Conclusion: some precise tests
Probability of distinctive results
(PDR)
Complete score comparison of pairs
Test I
Test II
Formula
sD
PDR
tD
A
B
C
Testwerte
A
B
C
Testwerte
Gaussian distribution
Rectangular distribution
shows a 60 %
shows an 80 %
probability to distinguish probability to distinguish
two test scores
two test scores
tD
n * (n 1)
2
n
s 1, if xi x j k
sD si , j i , j
si , j 0, if xi x j k
i, j
PDR: Simulation study
Performance to separate test scores with respect
to reliability and score distribution
PDR
Reliability
PDR: Example
SVF-KJ; Hampel, Petermann &
Dickow, 1999; N=1123
Subscale ‚Unsicherheit‘
Symptom Check List
(Derogatis, 1977; German
Version Franke, 1995; N=875
r = 0.81
r = 0.81
PDR = 41.6 %
PDR = 30.6 %
Subscale ‚Resignation‘;
Stress-Coping-Questionnaire
Reviewing NDR and PDR
1.
NDR and PDR can be derived in any test theoretical model –
there is progress in the application area.
2.
NDR and PDR have an easy to understand operational meaning
3.
NDR and PDR are predictable for the test user for the
nominal/actual value comparison
NDR and PDR serve as examples of how to develop more usable
formal test descriptions
Summary
1.
2.
Usability is a possible strategy with explicit and observable
criteria, for improving formal test descriptions – and
strengthening indirectly the role of guidelines and standards.
With NDR and PDR two easy to understood coefficients have
been proposed, the application of which in is progress in several
test theoretical models.
Thank you for your attention!
Medicine:
Effect-size measures
Scientific coefficient
(Cohen, 1988)
w
m
P1i P0i 2
i 1
P0i
Practitioners coefficient
1
NNT
RRR * CER
NNTs [Number-Needed-to-Treat] the
number of patients who need to be
treated to prevent 1 adverse
outcome.
Taken from EBM Glossary - Evidence
Based Medicine Volume 125 Number 1
Measuring in technical fields:
Solutions from engineering
The is a German Norm DIN 2257 on how to measure the
physical length of an object and how to report the result.
The norm allows as output only values with statistical
evidence.
Criteria of usability
for formal quality criteria for NNT
Foundation
1.
2.
3.
4.
5.
Unambiguous
operational
meaning
Unambiguous
formal definition
Broad application
area
Relevant
dependencies
Independent of
irrelevant factors
Scale Definition
1.
2.
3.
Meaningful scale unit,
that implies:
• Interval scale
• Positive values
• Defined range of
values
Comparable to the
reference scale
Significant scale unit, that
implies a minimum of
observations (Nmin)
Global attributes in
using
1.
2.
3.
4.
5.
6.
Relevance
Informative (not
redundant)
Predictable for the test
user (nominal/actual
value comparison)
Easy to learn
Easy to utilise
Fisher‘s (1925) criteria
of estimating
Criteria of software usability
(from Willumeit, Gediga & Hamborg, 1995)
Questionnaire on the basis of ISO9241/10 (IsoMetrics) to
evaluate the following dimensions:
1.
Suitability for the task
2.
Self-descriptiveness
3.
Controllability
4.
Conformity with user expectations
5.
Error tolerance
6.
Suitability for individualization
7.
Suitability for learning
KR20 and Cronbach
n
pi qi
Kuder-Richardson-Formel KR20
n u 1
i item
r
1
tt
pi relative Anzahl von 1
2
n 1
t
qi relative Anzahl von 0
(aus Cronbach, 1951)
Cronbachs Alpha
c Anzahl der Variablen
si2 Varianz der Variablen i
stot2 Varianz der Summe
J
2
si
c
1 i 1 2
c 1
sx
Formula to the error of measurement
in categorial constructs
A1
A2
B1
a
c
B2
b
d
ad
p0 pe
p
2
0
Cohen‘s Kappa
N
1 pe
Weiter 16 Maße zur Konkordanz zweier
(a c) (a b) (c d ) (b d )
Messungen für binaäre Daten verglichen
pe
Conger & Ward (1984)
N2
Yule
Vierfelderinterdependenzmaß
Q-Koeffizient
Phi-Koeffizient
Abhängigkeit von Randsummenverteilung
Abhängigkeit des Signifkanztests von N
(Yates-Kontinuitätskorrektur, 1934)
Y
1 bd ad
1 bc ad
ad bc
Q
ad bc
2
2
2
i 1 i 1
f
ij eij
2
eij
2
N
Formula to the error of measurement in
categorial constructs
Frickes Übereinstimmungskoeffizient SS: Quadratsumme innerhalb einer
Person; max SS: maximal mögliche Quadratsumme
innerhalb der Personen
Ü 1
SS
SSmax
ad
Ü
n
A
B
C
I
1
4
3
II
0
4
2
III 0
5
2
Punkt-biseriale Korrelation
X=arithmetisches Mittel aller Testrohwerte
XR=arithmetishes Mittel der Pbn mit richtigen Antworten
p _ bis rjt
sx=Standardabweichung der Testrohwerte aller Pbn
N = Anzahl aller Pbn
NR=Anzahl der Pbn, mit richtigen Antworten
Tetrachrorische Korrelation
XR X
sx
p
q
1800
rtet cos
1 ad bc
Formula to the error of measurement in CTT,
IRT + prophecy-formula
Spearman-Brown-Formel
k rtt
rtt
1 k 1rtt
k= Faktor der Testverlängerung
Rasch model
Var ( E )
1
k
p
i 1
vi
(1 pvi )
CTT
se sx 1 rtt
Some Formula for the error of measurement in
metric constructs
sw2
rtt 2 2
sw se
reliability (Kelley, 1921)
N
Pearson(1907) Correlation Bravais (1846)
Spearman‘s rho (1904)
Kendalls Tau , 1942
(S=difference of pro- und
inversionsnumber)
r
x x y y
i 1
i
i
N s x2 s y2
1
1 r
Z
ln
2
1 r
Rho 1
6 di2
i
N N 2 1
S
1
N ( N 1) / 2 3
2 3 4 5
2 3 5 4
Non-linear Relationsship between reliability,
NDR and the standard error score
reliability
NDR
1
NDR
Standard error score Standard error score
Item-Response-Theory
(Fischer & Molenaar, 1994)
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
Dichotomous raschmodel
Linear logistic test model
Linear logistic model for change
Dynamic generalization of the raschmodel
One parametric logistic model
Linear logistic latent class analysis
Mixture distribution rasch models
Polytomous rasch Models
Extended rating scale and partial credit models
Polytomous mixed rasch models
...
...more IRT
(van der Linden & Hambleton, 1997)
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
Nominal categories model
Response model for multiple choice
Graded response model
Partial credit model
Generalized partial credit model
Logistic model for time-limit tests
Hyperbolic cosine IRT model for unfolding direct responses
Single-item response model
Response model with manifest predictors
A linear multidimensional model
...
Formula of some IRT
binomial model
rasch model
Birnbaum model
Unfolding-model
Latent-Class-model
exp A
p x A
1 exp A
exp x Ai A i
p x Ai
1 exp A i
exp(xvi i ( v i ))
p( xvi )
1 exp( i ( v i ))
exp( xvi ( v i ) 2 )
p( xvi )
2
1 exp(( v i ) )
G
k
g 1
i 1
p( x) g igxi (1 ig )1 xi
Criteria of software usability
(from Willumeit, Gediga & Hamborg, 1995)
Questionnaire on the basis of ISO9241/10 (IsoMetrics) to
evaluate the following dimensions:
1.
Suitable for the task
2.
Self-descriptiveness
3.
Controllability
4.
Conformity with user expectations
5.
Error Tolerance
6.
Suitable for individualization
7.
Suitability for learning
Norm scales
SCL-90-R test score distribution
Simulation study about the relationsship
between measures of association
Normal distribution- equal marginals
dichotome
Measure of association
A1
A2
B1
a
c
B2
b
d
Skewed distribution - unequal marginals
Y/ Kappa/ Phi
Linear relationship?
correlation
Q
Measure of association
Y/ Kappa/ Phi
Y/ Kappa/ Phi
Phi
correlation
Y/Kappa/Phi
correlation
Q
SMC
Phi
SMC
SMC
Kappa
Kappa
Efficiency in measuring
Content:
Concept:
efficiency
The less effort you need for the same
amount of information, the more
efficiency the test is
efficiency = f(Information;effort)
Indice:
E = Amount of Information/Time
Estimates: Information Theory
(Shannon & Weaver, 1949)
Amount of Information of a signal:
Chess example
The scale unit ‚bit‘ can be understand as the minimal or
optimal number of question‘s, to identify a signal out of
quantity of alternatives.
1.Frage: linksrechts?
2.Frage: oben3.Frage
unten?
4.
5.
6.
In the chess example you
need at least (binary, 50-50
chance) 6 question‘s that
are 6 bit.
Rasch variances are a measure of the
variability of person‘s within a dimension
Schachspieler
B
1:2
1:2
1:2
A
1: 2
C
1: 2
1: 2
1: 2
1: 2Als
1: 2
Maßeinheit der
1: 2 Unterschiedlichkeit dient die
Differenz der
Gewinnwahrscheinlichkeiten.
Interpretable Rasch Variances
Probability to solve an item
1.
2.
3.
4.
Difference to solve a -> Lösungswahrscheinlichkeiten
Gewinnwahrscheinlichkeiten
Item i with = 0
question or task
Gegner -> Testaufgabe (Itemparameter)
A ->
max
min
Spielstärke
Personenparameter
item m with = 1
Differenz der Gewinnwahrscheinlichkeit definiert über den Logit
des Raschmodells
personen parameter
B
A
C
exp x Ai A i
p x Ai
1 exp A i
Empirical Evidence of the
range of person parameters in rasch units
AID Kubingen & Wurst
Standardform
Parallelform
Alltagswissen
21,1
21,3
Realitätssicherheit
13,3
13,1
Angewandtes Rechnen
21,7
20,5
Eigenschaft
Autor
Ausdehnung
Verbaler Intelligenztest
Averbale Intelligenz
Einstellung zur Sexualmoral
Einstellung zur Strafrechtsreform
Beschwerdeliste
Räumliches Vorstellungsvermögen
Umgang mit Zahlen bei Kindern
Metzler & Schmidt
Forman & Pieswanger
Wakenhut
Wakenhut
Fahrenberg
Gittler
Rasch
11,4
8,2
8,1
7,2
6,4
5,9
3,5
Usability criteria explanations
• Relevant dependencies: Example: Reliability and test length,
stability, ...
• Irrelevant dependencies: Example: Reliability and test score
distribution
• Displaying numbers: Integer, positive, predictable range
• Meaningful scale unit
• Familiarness: each new coefficient should distinctly more
usable than the traditional
7. Linearität zur Unit-in-Change
Erläuterung: ‚Linearität zur Unit-in-Change‘
- Im Falle der Messgenauigkeit betrifft dies die Beziehung der
Reliabilität zum Messfehler.
- Im Falle der Übereinstimmung betrifft dies die Beziehung von
Yules Y zur Veränderung der Zellhäufigkeit a bzw. d.
Korrelation/Reliabilität
Standardmessfehler
Yules Y
Freq (Zelle a)
Evaluation the progress trough
enhancing usability
1.
2.
Formal test criteria are used more
frequently for test selection
Tests in practice are of higher quality
Ergonomics in psychological test
selection
Ergonomics
Configuration of
Environment
Software
conception
Psychological
diagnostic
Designing a tool to
fit in hand.
Developing a
program to be used
intuitively
Restrict a test
description, that
relevant information
are ready to use
Integrating ergonomics in
the formal test description
Analysis
of usage
1.
2.
Usability
criteria
evaluation
Human
interface
techniques
test user
Psychometric
admeasurements
test
Formal test criteria are used more frequently for test selection
Tests in practice are of higher quality
Ergonomics and the development of
criteria of usability
Requirement Analysis (Mayhew,1999)
User-Profile
TaskAnalysis
Platform
Capabilities/
Constrains
Testuser
Test selection
Test theory
Top-down vs. bottom-up strategy
to develop a coefficient
Practitioner‘s point of view
association Interpretation of the score
(operational meaning)
P-R-E
none
N
r
i 1
scale (correction)
x x y y
i
i
N s x2 s y2
sw2
r 2
sw se2
CTT
Algorithm
Defining the operational
meaning
NDR
Scale definition
SEDTTS
Index: Defining the
influencing factors
Index: Generic formula
Index: Generic formula
test theory/statistic
Specification of within
a test theory
Scientist‘s point of
f(me, score range,
probability)
D
D
R
k
6 * sx
1,96 sx 21 rtt