CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Naiming, Baraniuk, Richard, Sonkar, Shashank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
by: Liu, Naiming, et al.
Published: (2025)
by: Liu, Naiming, et al.
Published: (2025)
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
by: Liu, Naiming, et al.
Published: (2024)
by: Liu, Naiming, et al.
Published: (2024)
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
LLM-based Cognitive Models of Students with Misconceptions
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Many-Shot Regurgitation (MSR) Prompting
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
by: Chen, Xinghe, et al.
Published: (2026)
by: Chen, Xinghe, et al.
Published: (2026)
Misconception Acquisition Dynamics in Large Language Models
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
Pedagogical Alignment of Large Language Models
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Circuit Complexity of Hierarchical Knowledge Tracing and Implications for Log-Precision Transformers
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
The Imitation Game for Educational AI
by: Sonkar, Shashank, et al.
Published: (2025)
by: Sonkar, Shashank, et al.
Published: (2025)
Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education
by: Liu, Naiming, et al.
Published: (2024)
by: Liu, Naiming, et al.
Published: (2024)
Synthetic Context Generation for Question Generation
by: Liu, Naiming, et al.
Published: (2024)
by: Liu, Naiming, et al.
Published: (2024)
Automated Long Answer Grading with RiceChem Dataset
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Scalable Generation and Validation of Isomorphic Physics Problems with GenAI
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
by: Ferrer, Robinson, et al.
Published: (2026)
by: Ferrer, Robinson, et al.
Published: (2026)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
by: Worden, Eamon, et al.
Published: (2026)
by: Worden, Eamon, et al.
Published: (2026)
CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
by: Huber, Thomas, et al.
Published: (2025)
by: Huber, Thomas, et al.
Published: (2025)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
by: Zengaffinen, Yanick, et al.
Published: (2026)
by: Zengaffinen, Yanick, et al.
Published: (2026)
LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
by: Chen, Junzhe, et al.
Published: (2024)
by: Chen, Junzhe, et al.
Published: (2024)
What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
by: Budding, Céline
Published: (2025)
by: Budding, Céline
Published: (2025)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
by: Masoud, Reem I., et al.
Published: (2023)
by: Masoud, Reem I., et al.
Published: (2023)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
by: Rufail, Andrew, et al.
Published: (2025)
by: Rufail, Andrew, et al.
Published: (2025)
CLEAR: Character Unlearning in Textual and Visual Modalities
by: Dontsov, Alexey, et al.
Published: (2024)
by: Dontsov, Alexey, et al.
Published: (2024)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
by: Tu, Ruibo, et al.
Published: (2024)
by: Tu, Ruibo, et al.
Published: (2024)
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
by: Shah, Kulin, et al.
Published: (2024)
by: Shah, Kulin, et al.
Published: (2024)
TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
by: Fujita, Aoi, et al.
Published: (2025)
by: Fujita, Aoi, et al.
Published: (2025)
CLEAR-KGQA: Clarification-Enhanced Ambiguity Resolution for Knowledge Graph Question Answering
by: Wen, Liqiang, et al.
Published: (2025)
by: Wen, Liqiang, et al.
Published: (2025)
Self-correction is Not An Innate Capability in Language Models
by: Liu, Guangliang, et al.
Published: (2024)
by: Liu, Guangliang, et al.
Published: (2024)
Sports Intelligence: Assessing the Sports Understanding Capabilities of Language Models through Question Answering from Text to Video
by: Yang, Zhengbang, et al.
Published: (2024)
by: Yang, Zhengbang, et al.
Published: (2024)
Assessing Generative Language Models in Classification Tasks: Performance and Self-Evaluation Capabilities in the Environmental and Climate Change Domain
by: Grasso, Francesca, et al.
Published: (2024)
by: Grasso, Francesca, et al.
Published: (2024)
CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
by: Lee, Seungyoon, et al.
Published: (2026)
by: Lee, Seungyoon, et al.
Published: (2026)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Explanatory Summarization with Discourse-Driven Planning
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
Similar Items
-
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
by: Liu, Naiming, et al.
Published: (2025) -
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning
by: Sonkar, Shashank, et al.
Published: (2024) -
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
by: Liu, Naiming, et al.
Published: (2024) -
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
by: Sonkar, Shashank, et al.
Published: (2024) -
MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support
by: Liu, Naiming, et al.
Published: (2026)