Abstract Item difficulty and dimensionality often correlate, implying that unidimensional IRT approximations to multidimensional data (i.e., reference composites) can take a curvilinear form in the multidimensional space. Although this issue has been previously discussed in the context of vertical scaling applications, we illustrate how such a phenomenon can also easily occur within individual tests. Measures of reading proficiency, for example, often use different task types within a single assessment, a feature that may not only lead to multidimensionality, but also an association between item difficulty and dimensionality. Using a latent regression strategy, we demonstrate through simulations and empirical analysis how associations between dimensionality and difficulty yield a nonlinear reference composite where the weights of the underlying dimensionschangeacross the scale continuum according to the difficulties of the items associated with the dimensions. We further show how this form of curvilinearity produces systematic forms of misspecification in traditional unidimensional IRT models (e.g., 2PL) and can be better accommodated by models such as monotone‐polynomial or asymmetric IRT models. Simulations and a real‐data example from the Early Childhood Longitudinal Study—Kindergarten are provided for demonstration. Some implications for measurement modeling and for understanding the effects of 2PL misspecification on measurement metrics are discussed.
more »
« less
This content will become publicly available on May 20, 2027
Approximating multidimensionality with asymmetric unidimensional IRT models
Unidimensional item response theory (IRT) models are widely used even in settings where assessment data exhibit subtle forms of multidimensionality. Recent empirical evidence suggests that when item difficulty is associated with dimensionality, asymmetric item characteristic curves (ICCs) emerge in the unidimensional approximation. Through theoretical derivation and extensive simulation, this paper develops a framework for understanding the emergence and the degree of ICC asymmetry in UIRT models when applied to multidimensional data with difficulty–dimensionality associations. An empirical analysis of the Virginia Language & Literacy Screener (VALLSS) confirms the predicted patterns. These results highlight ICC asymmetry as an anticipated consequence of unidimensional approximation and underscore its relevance for applications such as vertical scaling and test linking.
more »
« less
- Award ID(s):
- 2515523
- PAR ID:
- 10684202
- Publisher / Repository:
- Wiley, on behalf of the British Psychological Society
- Date Published:
- Journal Name:
- British Journal of Mathematical and Statistical Psychology
- ISSN:
- 0007-1102
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Fair and consistent assessment of student learning is critical in educational settings, particularly when evaluating the impact of instructional innovations. Although widely used for efficiency, output-based auto-grading often falls short in capturing partial understanding—limiting its effectiveness for measuring learning gains. This paper presents an empirical evaluation of a rubric-based, question-focused, double-grading protocol for written-response (WR) coding questions in pre- and post-tests from a large introductory programming course. This work provides both methodological insights and practical guidance for scaling reliable grading of WR coding questions. To balance efficiency and accuracy, each grader scored a specific question item across all submissions, with two graders assigned per item. Adjudication was triggered when score differences exceeded a 20% threshold. Intraclass Correlation Coefficient (ICC) analysis identified two items with initially low inter-rater reliability. After rubric clarification and regrading, reliability improved substantially, with ICC values ranging from 0.892 to 0.967 (all data) and 0.831 to 0.875 (excluding zero scores). We describe the iterative development of the assessment process and show how this structured approach—combined with ICC analysis as a diagnostic tool and targeted adjudication—achieves strong inter-grader reliability. The framework is scalable and robust for WR coding question evaluation in CS1 settings and is adaptable to a range of instructional contexts. These findings support instructors and researchers seeking consistent, practical methods for assessing WR student work in programming courses.more » « less
-
Deep reinforcement learning (RL) has shown remarkable success in specific offline decision-making scenarios, yet its theoretical guarantees are still under development. Existing works on offline RL theory primarily emphasize a few trivial settings, such as linear MDP or general function approximation with strong assumptions and independent data, which lack guidance for practical use. The coupling of deep learning and Bellman residuals makes this problem challenging, in addition to the difficulty of data dependence. In this paper, we establish a non-asymptotic estimation error of pessimistic offline RL using general neural network approximation with C-mixing data regarding the structure of networks, the dimension of datasets, and the concentrability of data coverage, under mild assumptions. Our result shows that the estimation error consists of two parts: the first converges to zero at a desired rate on the sample size with partially controllable concentrability, and the second becomes negligible if the residual constraint is tight. This result demonstrates the explicit efficiency of deep adversarial offline RL frameworks. We utilize the empirical process tool for C-mixing sequences and the neural network approximation theory for the Holder class to achieve this. We also develop methods to bound the Bellman estimation error caused by function approximation with empirical Bellman constraint perturbations. Additionally, we present a result that lessens the curse of dimensionality using data with low intrinsic dimensionality and function classes with low complexity. Our estimation provides valuable insights into the development of deep offline RL and guidance for algorithm model design.more » « less
-
Schwartz, Russell (Ed.)Abstract Summary Due to the sparsity and high dimensionality, microbiome data are routinely summarized into pairwise distances capturing the compositional differences. Many biological insights can be gained by analyzing the distance matrix in relation to some covariates. A microbiome sampling method that characterizes the inter-sample relationship more reproducibly is expected to yield higher statistical power. Traditionally, the intraclass correlation coefficient (ICC) has been used to quantify the degree of reproducibility for a univariate measurement using technical replicates. In this work, we extend the traditional ICC to distance measures and propose a distance-based ICC (dICC). We derive the asymptotic distribution of the sample-based dICC to facilitate statistical inference. We illustrate dICC using a real dataset from a metagenomic reproducibility study. Availability and implementation dICC is implemented in the R CRAN package GUniFrac. Supplementary information Supplementary data are available at Bioinformatics online.more » « less
-
While a common trend in disease modeling is to develop models of increasing complexity, it was recently pointed out that outbreaks appear remarkably simple when viewed in the incidence vs. cumulative cases (ICC) plane. This article details the theory behind this phenomenon by analyzing the stochastic Susceptible, Infected, Recovered (SIR) model in the cumulative cases domain. We prove that the Markov chain associated with this model reduces, in the ICC plane, to a pure birth chain for the cumulative number of cases, whose limit leads to an independent increments Gaussian process that fluctuates about a deterministic ICC curve. We calculate the associated variance and quantify the additional variability due to estimating incidence over a finite period of time. We also illustrate the universality brought forth by the ICC concept on real-world data for Influenza A and for the COVID-19 outbreak in Arizona.more » « less
An official website of the United States government
