Saved in:
| Main Authors: | Roemmele, Melissa, Gordon, Andrew S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.14897 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
by: Chung, John Joon Young, et al.
Published: (2025)
by: Chung, John Joon Young, et al.
Published: (2025)
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
by: Wang, Tiffany, et al.
Published: (2026)
by: Wang, Tiffany, et al.
Published: (2026)
Elsewise: Authoring AI-Based Interactive Narrative with Possibility Space Visualization
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
by: Colon-Hernandez, Pedro, et al.
Published: (2024)
by: Colon-Hernandez, Pedro, et al.
Published: (2024)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
by: Moore, Kyle, et al.
Published: (2024)
by: Moore, Kyle, et al.
Published: (2024)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
by: Ye, Yuxuan, et al.
Published: (2026)
by: Ye, Yuxuan, et al.
Published: (2026)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
by: Fish, Sara, et al.
Published: (2025)
by: Fish, Sara, et al.
Published: (2025)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
by: Nguyen, Bang, et al.
Published: (2025)
by: Nguyen, Bang, et al.
Published: (2025)
What Really is Commonsense Knowledge?
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
by: Ma, Yichuan, et al.
Published: (2026)
by: Ma, Yichuan, et al.
Published: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
Harnessing Consistency for Robust Test-Time LLM Ensemble
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
by: Stephan, Andreas, et al.
Published: (2024)
by: Stephan, Andreas, et al.
Published: (2024)
Learning From Mistakes Makes LLM Better Reasoner
by: An, Shengnan, et al.
Published: (2023)
by: An, Shengnan, et al.
Published: (2023)
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
by: Qiu, Ziqi, et al.
Published: (2024)
by: Qiu, Ziqi, et al.
Published: (2024)
Estimating Commonsense Plausibility through Semantic Shifts
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
From Data to Commonsense Reasoning: The Use of Large Language Models for Explainable AI
by: Krause, Stefanie, et al.
Published: (2024)
by: Krause, Stefanie, et al.
Published: (2024)
Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
by: Rosa, Bruno Alexandre, et al.
Published: (2025)
by: Rosa, Bruno Alexandre, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Modelling Commonsense Commonalities with Multi-Facet Concept Embeddings
by: Kteich, Hanane, et al.
Published: (2024)
by: Kteich, Hanane, et al.
Published: (2024)
Zero-shot Commonsense Reasoning over Machine Imagination
by: Park, Hyuntae, et al.
Published: (2024)
by: Park, Hyuntae, et al.
Published: (2024)
Commonsense Knowledge Editing Based on Free-Text in LLMs
by: Huang, Xiusheng, et al.
Published: (2024)
by: Huang, Xiusheng, et al.
Published: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
by: He, Mutian, et al.
Published: (2022)
by: He, Mutian, et al.
Published: (2022)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
by: Junias, Obed, et al.
Published: (2026)
by: Junias, Obed, et al.
Published: (2026)
Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models
by: Li, Chang-Jin, et al.
Published: (2024)
by: Li, Chang-Jin, et al.
Published: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024)
by: Bi, Zhenni, et al.
Published: (2024)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
by: Fang, Tianqing, et al.
Published: (2024)
by: Fang, Tianqing, et al.
Published: (2024)
Similar Items
-
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
by: Chung, John Joon Young, et al.
Published: (2025) -
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
by: Wang, Tiffany, et al.
Published: (2026) -
Elsewise: Authoring AI-Based Interactive Narrative with Possibility Space Visualization
by: Wang, Yi, et al.
Published: (2025) -
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
by: Colon-Hernandez, Pedro, et al.
Published: (2024) -
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)