From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Luoxi, Sundar, Tharunya, Meng, Yuqiao, Yang, Shuai, Patra, Ankita, Chippada, Lakshmi Manohar, Zhao, Jiqian, Li, Yi, Ma, Weicheng, Xi, Zhaohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
Adversarial Network Imagination: Causal LLMs and Digital Twins for Proactive Telecom Mitigation
by: Sriram, Vignesh, et al.
Published: (2026)
by: Sriram, Vignesh, et al.
Published: (2026)
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
Schema-Grounded LLM Extraction for FHIR Patient Digital Twins
by: Brens, Rafael, et al.
Published: (2026)
by: Brens, Rafael, et al.
Published: (2026)
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
by: Desai, Pratyush, et al.
Published: (2026)
by: Desai, Pratyush, et al.
Published: (2026)
Improving Clinical Data Accessibility Through Automated FHIR Data Transformation Tools
by: Pawar, Adarsh, et al.
Published: (2026)
by: Pawar, Adarsh, et al.
Published: (2026)
On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
RiskBridge: Turning CVEs into Business-Aligned Patch Priorities
by: Sheikh, Yelena Mujibur, et al.
Published: (2026)
by: Sheikh, Yelena Mujibur, et al.
Published: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
EquiMem: Calibrating Shared Memory in Multi-Agent Debate via Game-Theoretic Equilibrium
by: Meng, Yuqiao, et al.
Published: (2026)
by: Meng, Yuqiao, et al.
Published: (2026)
Small Agent Group is the Future of Digital Health
by: Meng, Yuqiao, et al.
Published: (2026)
by: Meng, Yuqiao, et al.
Published: (2026)
Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting
by: Liu, Xiaoqun, et al.
Published: (2025)
by: Liu, Xiaoqun, et al.
Published: (2025)
Diagnostic performance of online cognitive assessment and plasma biomarkers in Alzheimer’s Disease and Subjective Cognitive Impairment using the Oxford Cognitive Testing Portal
by: Sofia Toniolo, et al.
Published: (2024)
by: Sofia Toniolo, et al.
Published: (2024)
Diagnostic performance of online cognitive assessment and plasma biomarkers in Alzheimer’s Disease and Subjective Cognitive Impairment using the Oxford Cognitive Testing Portal
by: Sofia Toniolo, et al.
Published: (2024)
by: Sofia Toniolo, et al.
Published: (2024)
All Your Knowledge Belongs to Us: Stealing Knowledge Graphs via Reasoning APIs
by: Xi, Zhaohan
Published: (2025)
by: Xi, Zhaohan
Published: (2025)
Standardized Test Results: An Opportunity for English Program Improvement
by: Maureyra Jiménez
Published: (2017)
by: Maureyra Jiménez
Published: (2017)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Applying Cognitive Linguistics to Pedagogical Grammar: The English Prepositions of Verticality
by: Vyvyan Evans
Published: (2005)
by: Vyvyan Evans
Published: (2005)
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
by: Vico, Gianluca, et al.
Published: (2026)
by: Vico, Gianluca, et al.
Published: (2026)
A Forced-Choice Neural Cognitive Diagnostic Model of Personality Testing
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
English-Test International English Language Testing System PDF
by: Certification Exam
Published: (2026)
by: Certification Exam
Published: (2026)
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
by: Zhang, Zhaohan, et al.
Published: (2025)
by: Zhang, Zhaohan, et al.
Published: (2025)
Benchmarking Visual Language Models on Standardized Visualization Literacy Tests
by: Pandey, Saugat, et al.
Published: (2025)
by: Pandey, Saugat, et al.
Published: (2025)
Bootstrap Diagnostic Tests
by: Cavaliere, Giuseppe, et al.
Published: (2025)
by: Cavaliere, Giuseppe, et al.
Published: (2025)
Cash and Cognition: The Impact of Transfer Timing on Standardized Test Performance and Human Capital
by: Larrinaga, Axel Eizmendi, et al.
Published: (2025)
by: Larrinaga, Axel Eizmendi, et al.
Published: (2025)
An Assessment of Low-Power VLSI Testing Requires A Test Data Compression Architecture
by: P. Swetha, et al.
Published: (2021)
by: P. Swetha, et al.
Published: (2021)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
by: Momentè, Filippo, et al.
Published: (2025)
by: Momentè, Filippo, et al.
Published: (2025)
BLAST: Benchmarking LLMs with ASP-based Structured Testing
by: Santana, Manuel Alejandro Borroto, et al.
Published: (2026)
by: Santana, Manuel Alejandro Borroto, et al.
Published: (2026)
Evaluating the Test Adequacy of Benchmarks for LLMs on Code Generation
by: Xiangyue Liu, et al.
Published: (2025)
by: Xiangyue Liu, et al.
Published: (2025)
Effect of Imidazolium Concentration in Densely Functional Polymer Binder on Robust Electrochemical Kinetics and Cycling Performance of Lithium Iron Phosphate Cathode
by: Amarshi Patra, et al.
Published: (2025)
by: Amarshi Patra, et al.
Published: (2025)
GRE Practicing to take the biology test / Educational Testing Service
Published: (1995)
Published: (1995)
Estimating ordered variance of two scale mixture of normal distributions
by: Bajpai, Shrajal, et al.
Published: (2026)
by: Bajpai, Shrajal, et al.
Published: (2026)
Improved estimation of the positive powers ordered restricted standard deviation of two normal populations
by: Mondal, Somnath, et al.
Published: (2024)
by: Mondal, Somnath, et al.
Published: (2024)
Estimation of differential entropy for normal populations under prior information
by: Mandal, Somnath, et al.
Published: (2026)
by: Mandal, Somnath, et al.
Published: (2026)
Estimating order scale parameters of two scale mixture of exponential distributions
by: Mondal, Somnath, et al.
Published: (2025)
by: Mondal, Somnath, et al.
Published: (2025)
On the improved estimation of ordered parameters based on doubly type-II censored sample
by: Bajpai, Shrajal, et al.
Published: (2024)
by: Bajpai, Shrajal, et al.
Published: (2024)
Similar Items
-
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025) -
Adversarial Network Imagination: Causal LLMs and Digital Twins for Proactive Telecom Mitigation
by: Sriram, Vignesh, et al.
Published: (2026) -
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025) -
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
by: Kalvakuntla, Sriharshini, et al.
Published: (2026) -
Schema-Grounded LLM Extraction for FHIR Patient Digital Twins
by: Brens, Rafael, et al.
Published: (2026)