arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2607.27420v1 [cs.AI] 29 Jul 2026

Dimensionality and Measurement Precision in HLE’s Multiple-Choice Subset

Mayank Sharma
Stanford University
masharma@stanford.edu &Savira Nadela
Stanford University
savira@stanford.edu &Tyler Matteson
Stanford University
tylerjm@stanford.edu
Abstract

Humanity’s Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE (J=428J=428 items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald’s ωh=0.998\omega_{h}=0.998, domain labels explain only 3.5% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen’s d=0.016d=0.016), and domain-specific ability estimates are near-redundant with the total score (r0.81r\geq 0.81). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above θ=0\theta=0, where frontier models sit. These findings suggest that HLE’s domain subscores do not warrant distinct capability interpretations and that the benchmark’s ability to discriminate among the strongest models is limited.

1 Introduction

Humanity’s Last Exam (HLE) has quickly emerged as a prominent benchmark for evaluating advanced language models, appearing in capability assessments and policy discussions shortly after its release (Phan et al., 2026). Designed to resist simple retrieval and pattern matching through expert-level questions spanning mathematics, the natural sciences, and the humanities, HLE represents a plausible candidate for evaluating higher-order reasoning. However, widespread adoption has outpaced systematic evaluation of its measurement properties. Most benchmark studies, including HLE, report aggregate accuracy and domain-specific subscores without testing whether the reported domains correspond to empirically distinct latent factors. These subscores are often interpreted as evidence that one model outperforms another in domains such as mathematics or chemistry, an interpretation that implicitly assumes benchmark categories recover as separable latent capabilities rather than as facets of a single general reasoning factor (Jiang et al., 2026). Whether HLE’s domain labels reflect psychometrically distinct abilities remains largely untested, and the implications are concrete: if the eight domains do not correspond to distinct latent constructs, then per-domain model rankings reported in leaderboards are not warranted by the instrument, and developers and policy audiences should treat HLE as a measure of general reasoning rather than a profile of domain-specific skill.

A second concern is measurement precision. Even if HLE is unidimensional, its ability to differentiate between models depends on where along the ability continuum its items concentrate information. If most items are calibrated for average-performing models, score differences between the strongest frontier models may reflect measurement noise rather than genuine capability gaps, a problem that will intensify as models continue to improve.

This paper addresses both concerns directly. We evaluate 29 LLMs on the text-only MCQ subset of HLE (J=428J=428 items) and apply psychometric methods to investigate two questions: (a) Does HLE’s eight-domain structure reflect distinct latent constructs, or do the domains collapse into a general reasoning factor?, addressed through McDonald’s ωh\omega_{h}, PCA of item response profiles, residual correlation analysis, and domain-level ability comparisons; and (b) Where along the ability continuum does HLE concentrate measurement precision, and which domains contribute most to discrimination among frontier models?, addressed through the test information function decomposed by subject domain.

1.1 Related Work

Benchmarks are the primary instrument by which progress in large language models is evaluated, but their useful lifespan is often short. Popular benchmarks such as MMLU have shown signs of saturation within only a few years of release (Phan et al., 2026). This has motivated the development of HLE, a benchmark consisting of 2,500 expert-level questions spanning mathematics, humanities, and the natural sciences, crowdsourced from nearly 1,000 subject-matter experts across 50 countries and filtered to ensure that contemporary language models could not reliably answer them at submission time (Phan et al., 2026). HLE and related studies (e.g., MMLU (Hendrycks et al., 2021), GPQA (Rein et al., 2023)) typically report aggregate accuracy as the primary evaluation metric, but aggregate scores alone obscure whether items meaningfully discriminate between models and whether domain subscores reflect distinct latent constructs.

These measurement properties matter because benchmark rankings increasingly inform high-stakes decisions ranging from model deployment to AI governance (Liang et al., 2023). When domain subscores are interpreted as evidence of distinct capabilities, the validity of those comparisons is often assumed rather than empirically demonstrated, and decisions based on poorly validated measurements risk misrepresenting both relative model capability and the nature of progress in language model development (Bommasani et al., 2022).

Psychometric methods offer a framework for evaluating these properties directly. Applying IRT across 29 NLP datasets, Vania et al. (Vania et al., 2021) showed that several widely used benchmarks contain items that fail to discriminate effectively between models, raising questions about whether aggregate scores reflect meaningful capability differences. TinyBenchmarks (Polo et al., 2024) uses IRT to enable efficient evaluation on subsets of MMLU, and Chatbot Arena (Chiang et al., 2024) applies Bradley-Terry models to pairwise comparisons. These studies demonstrate the feasibility of treating benchmarks as measurement instruments, but they focus on ranking efficiency and item difficulty rather than on latent dimensionality: the question of whether domain-specific capability claims are structurally warranted. More broadly, benchmark leaderboards frequently report domain-specific performance as evidence of separable capabilities without testing whether such categories recover as empirically distinct latent factors (Jiang et al., 2026). To our knowledge, no psychometric dimensionality analysis has yet been applied to HLE. This study addresses that gap.

2 Methods

2.1 HLE Subset

The full benchmark contains two question formats: exact-match short-answer (\approx76%) and multiple-choice (\approx24%). Approximately 14% of all questions require image comprehension. Subject coverage skews toward quantitative domains: mathematics (41%), biology/medicine (11%), computer science/AI (10%), physics (9%), humanities/social science (9%), other (9%), chemistry (7%), and engineering (4%). We restrict to the text-only multiple-choice subset, which is the only subset amenable to automated binary scoring without multimodal infrastructure or LLM-judge extraction. This subset comprises J=513J=513 items, with individual items containing between 2 and 21 answer choices. The full dataset is publicly available via HuggingFace 111https://huggingface.co/datasets/cais/hle (cais/hle, split="test") and is filtered to this subset by selecting answer_type == "multipleChoice" and excluding image-dependent items. These items (see examples in Appendix A.2) demonstrate that they cannot be answered by retrieval or pattern matching and require domain expertise, making them valuable for an IRT analysis.

2.2 Model Selection

For generating responses on the subset, we used 37 contemporary language models across five model families: OpenAI (n=12n=12): GPT-4.1 series (standard, mini, nano), GPT-5 series (mini, nano, 5.4-mini, 5.4-nano), GPT-4o series (2024-11-20, mini), and reasoning models (o3-mini, o4-mini, o1-mini); Anthropic (n=6n=6): Claude Opus (4.7, 4.6, 4.5), Claude Sonnet (4.6, 4.5), and Claude Haiku 4.5; Google (n=5n=5): Gemini 3.5 Flash, Gemini 3.1 Flash Lite, Gemini 2.5 series (Pro, Flash, Flash Lite); DeepSeek (n=4n=4): DeepSeek-V3, DeepSeek-V4 (Flash, Pro), and DeepSeek-R1 (reasoning); Open-weight models (n=10n=10): Gemma (2-9B, 3-27B), Phi-4, Qwen2.5-7B, QwQ-32B (reasoning), Sky-T1-32B (reasoning), Mistral-7B, OLMo-2-7B, Falcon3-10B, and Llama-3.1-8B. Models represented both reasoning-specialized architectures and standard instruction-tuned models.

2.3 Response Collection

For each model-item pair, we prompted models using a standardized format (see prompt in Appendix A.1). To ensure reproducibility, we set temperature=0.0 for all models supporting this parameter. Reasoning-specialized models (OpenAI o-series, Claude Opus 4.7) used their default extended reasoning configurations without temperature control. Maximum response lengths were set to 8,192 tokens, with model-specific adjustments for known constraints (e.g., 4,096 tokens for smaller models). Proprietary models (OpenAI, Anthropic, Google, DeepSeek) were queried via their respective API endpoints with asynchronous request handling and automatic retry logic for rate-limit management. Open-weight models were deployed using vLLM (Kwon et al., 2023), a high-throughput inference engine optimized for large language models, on cloud GPU infrastructure (Modal Labs) using Nvidia A100-40GB, Nvidia A100-80GB, and Nvidia H100 GPUs. Data collection spanned approximately 20 hours, with the majority of latency attributable to reasoning-specialized models.

2.4 Response Parsing and Scoring

Model responses were parsed using rule-based extraction with a hierarchical strategy: (1) regex pattern matching for explicit answer declarations (“Answer:”, “Final Answer”), supporting markdown formatting and parenthetical notation ((A)); (2) flexible pattern matching for natural language phrasings (“the correct answer is A,” “choice is B”); and (3) fallback extraction of the last isolated letter (A-Z) when no explicit answer marker was present. Parsed responses (91.8%) were scored as correct (1) or incorrect (0) by exact-match comparison to the ground-truth answer. Unparsed responses (8.2%) were coded as missing data and predominantly resulted from empty model outputs (DeepSeek variants), safety-filtered refusals (Gemini 2.5 Pro), or inference failures (Phi-4, Gemma-9B).

2.5 Response Matrix Construction

We constructed a binary response matrix 𝐗{0,1,NA}N×J\mathbf{X}\in\{0,1,\text{NA}\}^{N\times J} where NN represents models and J=513J=513 items. From the initial 37 models evaluated, one model (o1-mini) was excluded due to zero coverage (0% valid responses). Coverage among the remaining 36 models ranged from 50.9% to 100%, with median 99.7% and mean 91.8%. Because missingness was not at random (attributable to content safety filtering, inference failures, and rate limiting rather than item difficulty), we retained only the 29 models with excellent coverage (\geq95%). Among the 513 items evaluated with these 29 models, 428 items (83.4%) exhibited non-zero variance across models; the remaining 85 items (16.6%) showed zero variance, with all models responding incorrectly, indicating extreme difficulty. We removed zero-variance items to obtain a final analytic matrix of 𝐗{0,1}29×428\mathbf{X}\in\{0,1\}^{29\times 428}, comprising 12,412 model-item observations, with remaining sparse missing values (0.79% of observations; n=117n=117) coded as incorrect responses, as their negligible proportion introduces minimal bias to estimates.

2.6 Analytic Overview

First, we ask whether HLE’s eight subject-domain labels correspond to empirically distinct latent constructs or collapse to a single general factor. Second, we ask how well HLE measures along the underlying factor(s): where measurement precision concentrates along the ability continuum, and which domains contribute most. We address both questions through a unified analysis, first estimating a two-parameter logistic (2PL) IRT model to obtain item-level discrimination and difficulty parameters, which serve as shared inputs to both stages, the dimensionality analysis and the measurement precision analysis.

2.6.1 2PL model estimation

We modeled item-level measurement properties using a two-parameter logistic (2PL) IRT model (Baker and Kim, 2004) fit to 𝐗\mathbf{X}. The probability that model ii correctly answers item jj was modeled as:

P(Xij=1θi,aj,bj)=11+exp[aj(θibj)],P(X_{ij}=1\mid\theta_{i},a_{j},b_{j})=\frac{1}{1+\exp[-a_{j}(\theta_{i}-b_{j})]}, (1)

where θi\theta_{i} represents latent model ability, aja_{j} is the item discrimination parameter, and bjb_{j} is the item difficulty parameter. Higher discrimination values indicate items that better differentiate between stronger and weaker models, whereas higher difficulty values indicate items requiring greater latent ability for a 50% probability of correct response. The model was estimated using marginal maximum likelihood implemented in torch_measure (Truong and others, 2026), with optimization run for up to 2,000 epochs using a learning rate of 0.05. For each item, we extracted estimated discrimination (a^j\hat{a}_{j}) and difficulty (b^j\hat{b}_{j}).

2.6.2 Dimensionality analysis

We examined whether the benchmark’s 8 subject-domain labels reflected empirically distinct latent constructs. Using data from N=29N=29 models and J=428J=428 items, we found that the inter-item tetrachoric correlation matrix was rank-deficient and non-positive-definite, preventing the use of confirmatory factor analysis because the resulting fit indices would be invalid (Flora and Curran, 2004). We therefore substituted three NN-robust alternatives that together address the same question without requiring inversion of the correlation matrix.

McDonald’s ωh\omega_{h}.

Our primary evidence for unidimensionality is McDonald’s hierarchical omega (McDonald, 1999), computed analytically from the 2PL discrimination parameters. Under the normal-ogive parameterization, item jj’s loading on the general factor is λj=aj/1+aj2\lambda_{j}=a_{j}/\sqrt{1+a_{j}^{2}}, and ωh\omega_{h} is defined as:

ωh=(jλj)2(jλj)2+j(1λj2)\omega_{h}=\frac{\left(\sum_{j}\lambda_{j}\right)^{2}}{\left(\sum_{j}\lambda_{j}\right)^{2}+\sum_{j}(1-\lambda_{j}^{2})} (2)

ωh\omega_{h} quantifies the proportion of item variance attributable to the general factor, ranging from 0 (no general factor) to 1 (perfectly unidimensional). Unlike CFA-based indices, this estimator requires no matrix inversion and is stable at small NN. We report ωh\omega_{h} overall and separately for each of the eight subject domains, with 95% bootstrap CIs constructed by resampling models with replacement (B=200B=200) and refitting the 2PL at each iteration.

PCA on item response profiles.

As a model-free complement, we transposed the response matrix to 𝐗{0,1}J×N\mathbf{X}^{\top}\in\{0,1\}^{J\times N}, treating each item as a point in NN-dimensional model space, and applied standard PCA. If domain labels capture real structure, items from the same domain should cluster together in this space. We quantified this using domain R2R^{2}, the proportion of variance in the first three principal components explained by domain membership, where R20R^{2}\approx 0 would indicate that items from the same domain respond no more similarly across models than items from different domains.

Residual item correlations.

To further assess whether domain membership explains structure beyond the general factor, we computed Pearson correlations between item response vectors, subtracted the general-factor-implied correlation matrix 𝐑g=𝝀𝝀\mathbf{R}_{g}=\boldsymbol{\lambda}\boldsymbol{\lambda}^{\top}, and compared within-domain to between-domain residual correlations. If domains capture distinct latent structure, within-domain residuals should be higher than between-domain residuals.

Domain-level θ^\hat{\theta} correlations.

To assess whether domain subscores carry any incremental information beyond the total score, we fit separate 2PL models within each domain and computed the Pearson and Spearman correlations between domain-specific ability estimates θ^domain\hat{\theta}_{\text{domain}} and overall ability estimates θ^\hat{\theta}. If domain subscores reflect distinct latent dimensions, models should exhibit differential ability profiles across domains; if a single underlying dimension dominates, r1r\approx 1 across all domains would indicate that the total score is sufficient and domain subscores carry little information.

2.6.3 Measurement precision

Taking the 2PL estimates and dimensionality findings together, we characterize where measurement precision concentrates along the ability continuum using the test information function (Baker and Kim, 2004),

I(θ)=j=1Jaj2Pj(θ)[1Pj(θ)],I(\theta)=\sum_{j=1}^{J}a_{j}^{2}P_{j}(\theta)\left[1-P_{j}(\theta)\right], (3)

which quantifies measurement precision at different ability levels, with higher values corresponding to lower conditional SEM. We decomposed the TIF by subject domain to evaluate which domains contributed most strongly to discrimination.

3 Results and Discussion

3.1 2PL Model Estimation

Across the full item pool, estimated difficulty parameters ranged from 2.06-2.06 to 5.675.67, with a median of b=0.41b=0.41, indicating that the benchmark was moderately challenging for the evaluated models. The positively skewed distribution further suggests a substantial concentration of advanced items capable of differentiating performance across a wide range of contemporary LLMs (Liang et al., 2023). Discrimination parameters ranged from 0.000.00 to the imposed upper cap of 5.005.00, with a median of a=1.49a=1.49. Many items demonstrated moderate-to-high discrimination values, indicating that the benchmark effectively differentiated between stronger and weaker models (see Figure 1). Mean empirical accuracy across items was relatively low (M=0.17M=0.17), further supporting the conclusion that the benchmark was challenging for most models. See Appendix A.3 for supplemental results.

Refer to caption
Figure 1: Distributions of item difficulty (bb) and item discrimination (aa).

Substantial variation emerged across disciplinary categories (see Figure 2). Engineering items exhibited the highest median difficulty (b=1.65b=1.65), followed by Physics (b=1.47b=1.47), whereas Computer Science/AI and Humanities/Social Science showed the lowest median difficulty values (both approximately b=0.10b=0.10). In contrast, discrimination patterns differed from difficulty trends. Computer Science/AI demonstrated the highest median discrimination (a=2.63a=2.63), while Engineering showed the lowest (a=0.77a=0.77), indicating that highly difficult domains did not necessarily provide the strongest differentiation between models. Category-level empirical accuracies aligned broadly with difficulty estimates, with Humanities/Social Science producing the highest mean accuracy (M=0.19M=0.19) and Engineering the lowest (M=0.13M=0.13). Figure 3 displays estimated latent ability θ^\hat{\theta} for all 29 models with 95% confidence intervals alongside raw accuracy. Claude Opus 4.7 achieved the highest estimated ability, followed by Gemini 3.5 Flash and Claude Opus 4.6. GPT-4o-2024-11-20 exhibited an unusually low ability estimate (θ^1.8\hat{\theta}\approx-1.8) with wide confidence intervals, suggesting unstable parameter estimation likely due to atypically low or inconsistent response patterns. IRT-based ability estimates were strongly correlated with raw accuracy rankings (ρ=0.857\rho=0.857, p<108p<10^{-8}), confirming that the 2PL model recovers a latent dimension consistent with aggregate performance while additionally characterizing item-level discrimination and measurement precision unavailable from raw scores alone. However, category-level differences in item parameters do not necessarily imply that HLE measures multiple distinct latent abilities: domains could still reflect a single underlying reasoning dimension if models that perform well in one domain tend to perform well across all others. The following analysis examines this directly.

Refer to caption
Figure 2: Category-level comparison of median item difficulty, median discrimination, and mean empirical accuracy across benchmark domains. Dashed vertical lines indicate overall median or mean values.
Refer to caption
Figure 3: Estimated ability θ^\hat{\theta} for all 29 models with 95% CIs (left axis), ordered by ability estimate, alongside raw accuracy (right axis, dashed). The two rankings are strongly correlated (ρ=0.857\rho=0.857).

3.2 Dimensionality Analysis

McDonald’s ωh\omega_{h}.

We found ωh=0.998\omega_{h}=0.998 (95% bootstrap CI [0.998,0.999][0.998,0.999]; B=200B=200 resamples), indicating that 99.8% of common item variance is attributable to a single general factor. The result is stable across all eight subject domains, with domain-level ωh\omega_{h} ranging from 0.9360.936 (Engineering, n=18n=18) to 0.9940.994 (Biology/Medicine, n=122n=122). The bootstrap SE of 0.00020.0002 confirms that this finding is not sensitive to the specific models in our analytic sample.

PCA on item response profiles.

Domain labels explained only 3.5%3.5\% of variance in the first three principal components (domain R2=0.035R^{2}=0.035), indicating that HLE’s subject categories impose no discernible structure on item response profiles beyond the general factor.

Residual item correlations.

Within-domain and between-domain residual correlations were nearly identical (μwithin=0.462\mu_{\text{within}}=-0.462, μbetween=0.466\mu_{\text{between}}=-0.466, Cohen’s d=0.016d=0.016). The near-zero effect size indicates that after removing the general factor, domain membership explains nothing about residual item correlation structure.222Absolute residual values are negative due to attenuation of Pearson rr on binary items; since attenuation affects within- and between-domain pairs equally, the comparison remains valid.

Domain-level θ^\hat{\theta} correlations.

For the four domains with sufficient item counts (n50n\geq 50), domain θ^\hat{\theta} was strongly correlated with overall θ^\hat{\theta} (Figure 4): Computer Science/AI (r=0.873r=0.873, ρ=0.928\rho=0.928), Biology/Medicine (r=0.866r=0.866, ρ=0.922\rho=0.922), Humanities/Social Science (r=0.859r=0.859, ρ=0.828\rho=0.828), and Mathematics (r=0.810r=0.810, ρ=0.816\rho=0.816; all p<107p<10^{-7}). Correlations for smaller domains (Engineering n=18n=18, Physics n=26n=26, Chemistry n=22n=22) were substantially attenuated and are not interpreted, as fitting a 2PL to fewer than 30 items with N=29N=29 models yields poorly identified parameter estimates. The near-identity relationship across all four well-powered domains indicates that a model’s total ability estimate is sufficient; domain subscores contribute little information beyond overall ability.

Refer to caption
Figure 4: θ^domain\hat{\theta}_{\text{domain}} vs. θ^\hat{\theta} for four domains (n50n\geq 50). Each point is a model, colored by domain accuracy. r0.81r\geq 0.81, ρ0.82\rho\geq 0.82, indicating near-redundancy between domain & total ability estimates.
Refer to caption
Figure 5: TIF decomposed by domain (stacked area). The shaded band marks the 10th-90th percentile ability range of models (θ[1.14,0.26]\theta\in[-1.14,-0.26]), showing that HLE concentrates measurement precision at moderate ability levels. The TIF drops sharply above θ=0\theta=0

3.3 Measurement Precision

The test information function (TIF) peaks at θ=0.35\theta=-0.35, coinciding with the ability range of most models in our sample (10th–90th percentile interval: [1.14,0.26][-1.14,-0.26]). Measurement precision drops substantially above θ=0\theta=0, precisely where the strongest-performing models sit. Decomposing the TIF by domain (Figure 5), Biology/Medicine contributes the most total information in absolute terms (I=12,611I=12{,}611). However, Engineering concentrates the highest proportion of its information at the frontier (69.1%69.1\% at θ>0\theta>0), consistent with its extreme item difficulty. Mathematics and Humanities/Social Science show the lowest frontier concentration (41.5%41.5\% and 41.4%41.4\% respectively), and despite containing more items than most other domains (n=78n=78 and n=71n=71), contribute disproportionately little to discrimination among higher-performing models.

4 Conclusion and Future Work

Fitting a 2PL IRT model to responses from 29 language models on HLE’s text-only multiple-choice subset (J=428J=428 items), we find convergent evidence that the benchmark measures a single general reasoning factor. McDonald’s ωh=0.998\omega_{h}=0.998 (95% CI [0.998,0.999][0.998,0.999]), domain labels explained only 3.5% of item response variance, residual correlations after removing the general factor were nearly identical within and between domains (Cohen’s d=0.016d=0.016), and domain-specific ability estimates were near-redundant with the total score (r0.81r\geq 0.81 across all well-powered domains). HLE’s eight subject-domain labels do not correspond to empirically distinct latent constructs; so they might be arbitrary partitions of a unidimensional space. A separate finding concerns measurement precision. While HLE discriminates well among models at moderate ability levels, precision drops substantially above θ=0\theta=0, precisely where frontier models sit. Engineering items concentrate the most information at the frontier (69.1%69.1\% at θ>0\theta>0) but represent only 4% of the benchmark; Mathematics and Humanities/Social Science, the two largest domains, contribute disproportionately little to frontier discrimination. As stronger models continue to emerge, this precision gap will become an increasingly binding constraint on HLE’s utility as a differentiating instrument.

The psychometric approach worked well for this setting: IRT discrimination parameters proved stable enough to support downstream dimensionality analyses, and the convergence of four independent analyses (ωh\omega_{h}, domain R2R^{2}, residual correlations, and θ^\hat{\theta} redundancy) substantially strengthens confidence in the unidimensionality conclusion. The primary methodological challenge was the small model sample (N=29N=29), which precluded confirmatory factor analysis and limited statistical power for domain-level analyses. This inverted the usual psychometric setting, where items are typically far fewer than subjects, and required N-robust alternatives throughout. A practical lesson for future benchmark psychometrics is that model coverage must be treated as a first-class design consideration: missingness that is not at random, as observed here for reasoning-specialized models, systematically biases the analytic sample and limits generalizability.

Several limitations bound these conclusions. Our analytic sample is restricted to the text-only multiple-choice subset (\approx19% of HLE), so findings do not extend to exact-match or image-dependent items; moreover, the multiple-choice format may introduce construct-irrelevant variance if models succeed through elimination rather than genuine domain knowledge. The model sample is small (N=29N=29) and non-random, as eight models were excluded for low response coverage, disproportionately affecting reasoning-specialized architectures and precluding formal measurement invariance testing across model families. Scoring each model on a single deterministic response per item means discrimination estimates reflect stable model differences but not within-model response variability, which may underestimate uncertainty in item parameter estimates. Future work should replicate on the full benchmark with a larger and more complete model sample to enable CFA and invariance testing.

Code Availability

All analysis code used in this study is available at: https://github.com/matrix-mayank/hle-psychometrics.

AI Usage

AI tools were used in a limited way throughout this project, mainly to help improve writing clarity, troubleshoot coding issues, and brainstorm ways to present results more clearly. To avoid plagiarism or non-attribution of ideas, all citations and references were added manually by the authors after checking the original academic sources directly. We also reviewed AI-generated suggestions carefully to make sure ideas, interpretations, and wording were appropriately credited and not copied without attribution. Responsibility for the originality, accuracy, and scientific integrity of the final paper remained entirely with the authors.

Ethics Statement

This project examines the psychometric properties of a large language model benchmark and highlights the importance of carefully interpreting benchmark scores and domain subscores. One potential positive impact is encouraging more rigorous evaluation of benchmark validity, especially when benchmark rankings are used in research, industry, or policy discussions. However, the findings could also be misinterpreted if generalized beyond the analyzed subset or treated as definitive statements about model intelligence. To reduce this risk, the paper clearly discusses limitations such as the small model sample and restriction to the multiple-choice subset of HLE.

References

  • F. B. Baker and S. Kim (2004) Item response theory: parameter estimation techniques. 2 edition, CRC Press. External Links: Document Cited by: §2.6.1, §2.6.3.
  • R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang (2022) On the opportunities and risks of foundation models. External Links: 2108.07258, Link Cited by: §1.1.
  • W. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. I. Jordan, J. E. Gonzalez, and I. Stoica (2024) Chatbot arena: an open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132. External Links: Link Cited by: §1.1.
  • D. B. Flora and P. J. Curran (2004) An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data. Psychological Methods 9 (4), pp. 466–491. External Links: Document Cited by: §2.6.2.
  • D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021) Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300. External Links: 2009.03300, Link Cited by: §1.1.
  • H. Jiang, S. Zhang, D. Zhu, Y. Bai, S. T. Truong, X. Yi, S. Koyejo, X. Xie, and Z. Xiao (2026) AI evaluation should require standardized item-level data releases. arXiv preprint arXiv:2604.03244. External Links: Link Cited by: §1.1, §1.
  • W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica (2023) Efficient memory management for large language model serving with pagedattention. External Links: 2309.06180, Link Cited by: §2.3.
  • P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al. (2023) Holistic evaluation of language models. Transactions on Machine Learning Research. External Links: 2211.09110, Link Cited by: §1.1, §3.1.
  • R. P. McDonald (1999) Test theory: a unified treatment. 1 edition, Psychology Press. External Links: Document Cited by: §2.6.2.
  • L. Phan, A. Gatti, N. Li, A. Khoja, R. Kim, R. Ren, J. Hausenloy, O. Zhang, M. Mazeika, D. Hendrycks, Z. Han, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, M. Choi, A. Agrawal, A. Chopra, A. Nattanmai, G. McKellips, A. Cheraku, A. Suhail, E. Luo, M. Deng, J. Luo, A. Zhang, K. Jindel, J. Paek, K. Halevy, A. Baranov, M. Liu, A. Avadhanam, D. Zhang, V. Cheng, B. Ma, E. Fu, L. Do, J. Lass, H. Yang, S. Sunkari, V. Bharath, V. Ai, J. Leung, R. Agrawal, A. Zhou, K. Chen, T. Kalpathi, Z. Xu, G. Wang, T. Xiao, E. Maung, S. Lee, R. Yang, R. Yue, B. Zhao, J. Yoon, X. Sun, A. Singh, C. Peng, T. Osbey, T. Wang, D. Echeazu, T. Wu, S. Patel, V. Kulkarni, V. Sundarapandiyan, A. Le, Z. Nasim, S. Yalam, R. Kasamsetty, S. Samal, D. Sun, N. Shah, A. Saha, A. Zhang, L. Nguyen, L. Nagumalli, K. Wang, A. Wu, A. Telluri, S. Yue, A. Wang, D. Dodonov, T. Nguyen, J. Lee, D. Anderson, M. Doroshenko, A. C. Stokes, M. Mahmood, O. Pokutnyi, O. Iskra, J. P. Wang, J. Levin, M. Kazakov, F. Feng, S. Y. Feng, H. Zhao, M. Yu, V. Gangal, C. Zou, Z. Wang, S. Popov, R. Gerbicz, G. Galgon, J. Schmitt, W. Yeadon, Y. Lee, S. Sauers, A. Sanchez, F. Giska, M. Roth, S. Riis, S. Utpala, N. Burns, G. M. Goshu, M. M. Naiya, C. Agu, Z. Giboney, A. Cheatom, F. Fournier-Facio, S. Crowson, L. Finke, Z. Cheng, J. Zampese, R. G. Hoerr, M. Nandor, H. Park, T. Gehrunger, J. Cai, B. McCarty, A. C. Garretson, E. Taylor, D. Sileo, Q. Ren, U. Qazi, L. Li, J. Nam, J. B. Wydallis, P. Arkhipov, J. W. L. Shi, A. Bacho, C. G. Willcocks, H. Cao, S. Motwani, E. de Oliveira Santos, J. Veith, E. Vendrow, D. Cojoc, K. Zenitani, J. Robinson, L. Tang, Y. Li, J. Vendrow, N. W. Fraga, V. Kuchkin, A. P. Maksimov, P. Marion, D. Efremov, J. Lynch, K. Liang, A. Mikov, A. Gritsevskiy, J. Guillod, G. Demir, D. Martinez, B. Pageler, K. Zhou, S. Soori, O. Press, H. Tang, P. Rissone, S. R. Green, L. Brüssel, M. Twayana, A. Dieuleveut, J. M. Imperial, A. Prabhu, J. Yang, N. Crispino, A. Rao, D. Zvonkine, G. Loiseau, M. Kalinin, M. Lukas, C. Manolescu, N. Stambaugh, S. Mishra, T. Hogg, C. Bosio, B. P. Coppola, J. Salazar, J. Jin, R. Sayous, S. Ivanov, P. Schwaller, S. Senthilkumar, A. M. Bran, A. Algaba, K. Van den Houte, L. Van Der Sypt, B. Verbeken, D. Noever, A. Kopylov, B. Myklebust, B. Li, L. Schut, E. Zheltonozhskii, Q. Yuan, D. Lim, R. Stanley, T. Yang, J. Maar, J. Wykowski, M. Oller, A. Sahu, C. G. Ardito, Y. Hu, A. G. K. Kamdoum, A. Jin, T. G. Vilchis, Y. Zu, M. Lackner, J. Koppel, G. Sun, D. S. Antonenko, S. Chern, B. Zhao, P. Arsene, J. M. Cavanagh, D. Li, J. Shen, D. Crisostomi, W. Zhang, A. Dehghan, S. Ivanov, D. Perrella, N. Kaparov, A. Zang, I. Sucholutsky, A. Kharlamova, D. Orel, V. Poritski, S. Ben-David, Z. Berger, P. Whitfill, M. Foster, D. Munro, L. Ho, S. Sivarajan, D. B. Hava, A. Kuchkin, D. Holmes, A. Rodriguez-Romero, F. Sommerhage, A. Zhang, R. Moat, K. Schneider, Z. Kazibwe, D. Clarke, D. H. Kim, F. M. Dias, S. Fish, V. Elser, T. Kreiman, V. E. G. Vilchis, I. Klose, U. Anantheswaran, A. Zweiger, K. Rawal, J. Li, J. Nguyen, N. Daans, H. Heidinger, M. Radionov, V. Rozhoň, V. Ginis, C. Stump, N. Cohen, R. Poświata, J. Tkadlec, A. Goldfarb, C. Wang, P. Padlewski, S. Barzowski, K. Montgomery, R. Stendall, J. Tucker-Foltz, J. Stade, T. R. Rogers, T. Goertzen, D. Grabb, A. Shukla, A. Givré, J. A. Ambay, A. Sen, Center for AI Safety, Scale AI, and HLE Contributors Consortium (2026) A benchmark of expert-level academic questions to assess ai capabilities. Nature 649 (8099), pp. 1139–1146. External Links: Document, Link, ISSN 1476-4687 Cited by: §1.1, §1.
  • F. M. Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin (2024) TinyBenchmarks: evaluating llms with fewer examples. arXiv preprint arXiv:2402.14992. External Links: Link Cited by: §1.1.
  • D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman (2023) GPQA: a graduate-level google-proof q&a benchmark. External Links: 2311.12022, Link Cited by: §1.1.
  • S. T. Truong et al. (2026) Torch_measure: a package for ai measurement science Note: MIT License External Links: Link Cited by: §2.6.1.
  • C. Vania, P. M. Htut, W. Huang, D. Mungra, R. Y. Pang, J. Phang, H. Liu, K. Cho, and S. R. Bowman (2021) Comparing test sets with item response theory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, pp. 1141–1158. External Links: Document, Link Cited by: §1.1.

Appendix A Appendix

A.1 Response Collection Prompt

Standardized Prompt Format Your response should be in the following format: Explanation: {your explanation for your answer choice} Answer: {your chosen answer} Confidence: {your confidence score between 0% and 100% for your answer}

A.2 Example Items from HLE

Example Items from HLE 1. Philosophy: “Which condition of Arrhenius’s sixth impossibility theorem do critical-level views violate?” with answer choices including Egalitarian Dominance, General Non-Extreme Priority, Non-Elitism, Weak Non-Sadism, and Weak Quality Addition. 2. Linguistics: “What is the standard Japanese pitch accent pattern of the word for ‘younger brother’ in Japanese?” with choices spanning Heiban, Atamadaka, Nakadaka, Odaka, and Heiban/Nakadaka.

A.3 Additional 2PL Diagnostic Plots

Figure 6 presents supplementary diagnostic visualizations for the estimated 2PL parameters. The left panel illustrates the inverse relationship between item difficulty and discrimination, where extremely difficult items tended to exhibit lower discrimination values. The right panel shows the expected negative relationship between empirical accuracy and estimated item difficulty, supporting the consistency of the estimated IRT parameters.

Refer to caption
Figure 6: Supplementary 2PL diagnostic plots. Left: relationship between item difficulty and discrimination across benchmark domains. Right: empirical item accuracy against estimated 2PL difficulty.