From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Roemmele, Melissa, Gordon, Andrew S. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
par: Chung, John Joon Young, et autres
Publié: (2025)
par: Chung, John Joon Young, et autres
Publié: (2025)
Self-Assessment Tests are Unreliable Measures of LLM Personality
par: Gupta, Akshat, et autres
Publié: (2023)
par: Gupta, Akshat, et autres
Publié: (2023)
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
par: Colon-Hernandez, Pedro, et autres
Publié: (2024)
par: Colon-Hernandez, Pedro, et autres
Publié: (2024)
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
par: Wang, Tiffany, et autres
Publié: (2026)
par: Wang, Tiffany, et autres
Publié: (2026)
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
par: Moore, Kyle, et autres
Publié: (2024)
par: Moore, Kyle, et autres
Publié: (2024)
Elsewise: Authoring AI-Based Interactive Narrative with Possibility Space Visualization
par: Wang, Yi, et autres
Publié: (2025)
par: Wang, Yi, et autres
Publié: (2025)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
par: Ye, Yuxuan, et autres
Publié: (2026)
par: Ye, Yuxuan, et autres
Publié: (2026)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
par: Cheng, Qi, et autres
Publié: (2024)
par: Cheng, Qi, et autres
Publié: (2024)
What Really is Commonsense Knowledge?
par: Do, Quyet V., et autres
Publié: (2024)
par: Do, Quyet V., et autres
Publié: (2024)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
par: Fish, Sara, et autres
Publié: (2025)
par: Fish, Sara, et autres
Publié: (2025)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
par: Ma, Yichuan, et autres
Publié: (2026)
par: Ma, Yichuan, et autres
Publié: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2024)
par: Kim, Eunsu, et autres
Publié: (2024)
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
par: Nguyen, Bang, et autres
Publié: (2025)
par: Nguyen, Bang, et autres
Publié: (2025)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
par: Stephan, Andreas, et autres
Publié: (2024)
par: Stephan, Andreas, et autres
Publié: (2024)
Benchmark Test-Time Scaling of General LLM Agents
par: Li, Xiaochuan, et autres
Publié: (2026)
par: Li, Xiaochuan, et autres
Publié: (2026)
Harnessing Consistency for Robust Test-Time LLM Ensemble
par: Zeng, Zhichen, et autres
Publié: (2025)
par: Zeng, Zhichen, et autres
Publié: (2025)
Learning From Mistakes Makes LLM Better Reasoner
par: An, Shengnan, et autres
Publié: (2023)
par: An, Shengnan, et autres
Publié: (2023)
From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
par: Tang, Luoxi, et autres
Publié: (2025)
par: Tang, Luoxi, et autres
Publié: (2025)
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
par: Qiu, Ziqi, et autres
Publié: (2024)
par: Qiu, Ziqi, et autres
Publié: (2024)
Estimating Commonsense Plausibility through Semantic Shifts
par: Cui, Wanqing, et autres
Publié: (2025)
par: Cui, Wanqing, et autres
Publié: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
par: Balkır, Esma, et autres
Publié: (2026)
par: Balkır, Esma, et autres
Publié: (2026)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
par: Bi, Zhenni, et autres
Publié: (2024)
par: Bi, Zhenni, et autres
Publié: (2024)
Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
par: Rosa, Bruno Alexandre, et autres
Publié: (2025)
par: Rosa, Bruno Alexandre, et autres
Publié: (2025)
Modelling Commonsense Commonalities with Multi-Facet Concept Embeddings
par: Kteich, Hanane, et autres
Publié: (2024)
par: Kteich, Hanane, et autres
Publié: (2024)
Zero-shot Commonsense Reasoning over Machine Imagination
par: Park, Hyuntae, et autres
Publié: (2024)
par: Park, Hyuntae, et autres
Publié: (2024)
Commonsense Knowledge Editing Based on Free-Text in LLMs
par: Huang, Xiusheng, et autres
Publié: (2024)
par: Huang, Xiusheng, et autres
Publié: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
par: Yang, Shuo, et autres
Publié: (2024)
par: Yang, Shuo, et autres
Publié: (2024)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
par: Wang, Jingxing, et autres
Publié: (2026)
par: Wang, Jingxing, et autres
Publié: (2026)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
par: Yang, Wenkai, et autres
Publié: (2025)
par: Yang, Wenkai, et autres
Publié: (2025)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
par: Li, Peiyu, et autres
Publié: (2025)
par: Li, Peiyu, et autres
Publié: (2025)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
par: Yang, Zhen, et autres
Publié: (2025)
par: Yang, Zhen, et autres
Publié: (2025)
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
par: He, Mutian, et autres
Publié: (2022)
par: He, Mutian, et autres
Publié: (2022)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
par: Junias, Obed, et autres
Publié: (2026)
par: Junias, Obed, et autres
Publié: (2026)
From Data to Commonsense Reasoning: The Use of Large Language Models for Explainable AI
par: Krause, Stefanie, et autres
Publié: (2024)
par: Krause, Stefanie, et autres
Publié: (2024)
From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness
par: Cao, Linbo, et autres
Publié: (2026)
par: Cao, Linbo, et autres
Publié: (2026)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
par: Liu, Yunting, et autres
Publié: (2024)
par: Liu, Yunting, et autres
Publié: (2024)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
par: Tang, Zeyu, et autres
Publié: (2026)
par: Tang, Zeyu, et autres
Publié: (2026)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
par: Fang, Tianqing, et autres
Publié: (2024)
par: Fang, Tianqing, et autres
Publié: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
par: Palta, Shramay, et autres
Publié: (2024)
par: Palta, Shramay, et autres
Publié: (2024)
Self-Improving LLM Agents at Test-Time
par: Acikgoz, Emre Can, et autres
Publié: (2025)
par: Acikgoz, Emre Can, et autres
Publié: (2025)
Documents similaires
-
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
par: Chung, John Joon Young, et autres
Publié: (2025) -
Self-Assessment Tests are Unreliable Measures of LLM Personality
par: Gupta, Akshat, et autres
Publié: (2023) -
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
par: Colon-Hernandez, Pedro, et autres
Publié: (2024) -
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
par: Wang, Tiffany, et autres
Publié: (2026) -
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
par: Moore, Kyle, et autres
Publié: (2024)