JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Louie Hong, Jarvis, Nicholas, Zhan, Tiffany, Ghosh, Saptarshi, Liu, Linfeng, Jiang, Tianyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
Exploring Concreteness Through a Figurative Lens
by: Ghosh, Saptarshi, et al.
Published: (2026)
by: Ghosh, Saptarshi, et al.
Published: (2026)
Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
by: Liu, Linfeng, et al.
Published: (2025)
by: Liu, Linfeng, et al.
Published: (2025)
A Computational Approach to Visual Metonymy
by: Ghosh, Saptarshi, et al.
Published: (2026)
by: Ghosh, Saptarshi, et al.
Published: (2026)
MetFuse: Figurative Fusion between Metonymy and Metaphor
by: Ghosh, Saptarshi, et al.
Published: (2026)
by: Ghosh, Saptarshi, et al.
Published: (2026)
ConMeC: A Dataset for Metonymy Resolution with Common Nouns
by: Ghosh, Saptarshi, et al.
Published: (2025)
by: Ghosh, Saptarshi, et al.
Published: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
by: Yao, Louie Hong, et al.
Published: (2026)
by: Yao, Louie Hong, et al.
Published: (2026)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
by: Uebayashi, Shunki, et al.
Published: (2026)
by: Uebayashi, Shunki, et al.
Published: (2026)
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
by: Vennemeyer, Daniel, et al.
Published: (2025)
by: Vennemeyer, Daniel, et al.
Published: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series
by: Sun, Chenxi, et al.
Published: (2023)
by: Sun, Chenxi, et al.
Published: (2023)
AutoIRT: Calibrating Item Response Theory Models with Automated Machine Learning
by: Sharpnack, James, et al.
Published: (2024)
by: Sharpnack, James, et al.
Published: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Enhancing Biochemistry Assessment Quality in Medical Education Through Item Response Theory ( IRT )
by: Baharuddin Baharuddin, et al.
Published: (2025)
by: Baharuddin Baharuddin, et al.
Published: (2025)
Learning Compact Representations of LLM Abilities via Item Response Theory
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
by: Li, Xinyuan, et al.
Published: (2025)
by: Li, Xinyuan, et al.
Published: (2025)
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
by: Zhou, Hongli, et al.
Published: (2025)
by: Zhou, Hongli, et al.
Published: (2025)
Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
by: Dhanuka, Utsav, et al.
Published: (2025)
by: Dhanuka, Utsav, et al.
Published: (2025)
LLM Bias Detection and Mitigation through the Lens of Desired Distributions
by: Shrestha, Ingroj, et al.
Published: (2025)
by: Shrestha, Ingroj, et al.
Published: (2025)
Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
by: Rosa, Bruno Alexandre, et al.
Published: (2025)
by: Rosa, Bruno Alexandre, et al.
Published: (2025)
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024)
by: Schwarzschild, Avi, et al.
Published: (2024)
Synthesizing the Ability in Multidimensional Item Response Theory Models
by: ÁLVARO MAURICIO MONTENEGRO DÍAZ
Published: (2010)
by: ÁLVARO MAURICIO MONTENEGRO DÍAZ
Published: (2010)
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
by: Lin, Fan, et al.
Published: (2024)
by: Lin, Fan, et al.
Published: (2024)
Improve LLM-as-a-Judge Ability as a General Ability
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Modeling Abilities in 3-IRT Models
by: Edilberto Cepeda
Published: (2004)
by: Edilberto Cepeda
Published: (2004)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Do LLM hallucination detectors suffer from low-resource effect?
by: Datta, Debtanu, et al.
Published: (2026)
by: Datta, Debtanu, et al.
Published: (2026)
Automated Essay Scoring Using Grammatical Variety and Errors with Multi-Task Learning and Item Response Theory
by: Doi, Kosuke, et al.
Published: (2024)
by: Doi, Kosuke, et al.
Published: (2024)
Error as a Lens: Probing LLM Reasoning through Synthetic Misconception Generation
by: Yang, Xinming, et al.
Published: (2026)
by: Yang, Xinming, et al.
Published: (2026)
Beyond Borders: Investigating Cross-Jurisdiction Transfer in Legal Case Summarization
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
ICPR 2024 Competition on Multilingual Claim-Span Identification
by: Poddar, Soham, et al.
Published: (2024)
by: Poddar, Soham, et al.
Published: (2024)
ILSIC: Corpora for Identifying Indian Legal Statutes from Queries by Laypeople
by: Paul, Shounak, et al.
Published: (2026)
by: Paul, Shounak, et al.
Published: (2026)
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
Brevity is the soul of sustainability: Characterizing LLM response lengths
by: Poddar, Soham, et al.
Published: (2025)
by: Poddar, Soham, et al.
Published: (2025)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
MAQuA: Adaptive Question-Asking for Multidimensional Mental Health Screening using Item Response Theory
by: Varadarajan, Vasudha, et al.
Published: (2025)
by: Varadarajan, Vasudha, et al.
Published: (2025)
Similar Items
-
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025) -
Exploring Concreteness Through a Figurative Lens
by: Ghosh, Saptarshi, et al.
Published: (2026) -
Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
by: Liu, Linfeng, et al.
Published: (2025) -
A Computational Approach to Visual Metonymy
by: Ghosh, Saptarshi, et al.
Published: (2026) -
MetFuse: Figurative Fusion between Metonymy and Metaphor
by: Ghosh, Saptarshi, et al.
Published: (2026)