Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Ahuja, Kabir, Sclar, Melanie, Tsvetkov, Yulia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023)
by: Sclar, Melanie, et al.
Published: (2023)
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
by: Faisal, Fahim, et al.
Published: (2024)
by: Faisal, Fahim, et al.
Published: (2024)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
by: Sclar, Melanie, et al.
Published: (2024)
by: Sclar, Melanie, et al.
Published: (2024)
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
by: Ahuja, Kabir, et al.
Published: (2024)
by: Ahuja, Kabir, et al.
Published: (2024)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
by: Wang, Yike, et al.
Published: (2025)
by: Wang, Yike, et al.
Published: (2025)
Fine-grained Hallucination Detection and Editing for Language Models
by: Mishra, Abhika, et al.
Published: (2024)
by: Mishra, Abhika, et al.
Published: (2024)
KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models
by: Bai, Yuyang, et al.
Published: (2023)
by: Bai, Yuyang, et al.
Published: (2023)
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Knowledge Crosswords: Geometric Knowledge Reasoning with Large Language Models
by: Ding, Wenxuan, et al.
Published: (2023)
by: Ding, Wenxuan, et al.
Published: (2023)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
by: Jiang, Yuru, et al.
Published: (2025)
by: Jiang, Yuru, et al.
Published: (2025)
MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
by: Wang, Haojin, et al.
Published: (2026)
by: Wang, Haojin, et al.
Published: (2026)
The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems
by: Feng, Shangbin, et al.
Published: (2026)
by: Feng, Shangbin, et al.
Published: (2026)
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
by: Yang, Ziyuan, et al.
Published: (2026)
by: Yang, Ziyuan, et al.
Published: (2026)
Tuning Language Models by Proxy
by: Liu, Alisa, et al.
Published: (2024)
by: Liu, Alisa, et al.
Published: (2024)
ComPO: Community Preferences for Language Model Personalization
by: Kumar, Sachin, et al.
Published: (2024)
by: Kumar, Sachin, et al.
Published: (2024)
Small Reward Models via Backward Inference
by: Wang, Yike, et al.
Published: (2026)
by: Wang, Yike, et al.
Published: (2026)
Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
by: Feng, Shangbin, et al.
Published: (2023)
by: Feng, Shangbin, et al.
Published: (2023)
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Can Language Models Solve Graph Problems in Natural Language?
by: Wang, Heng, et al.
Published: (2023)
by: Wang, Heng, et al.
Published: (2023)
Resolving Knowledge Conflicts in Large Language Models
by: Wang, Yike, et al.
Published: (2023)
by: Wang, Yike, et al.
Published: (2023)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
by: Qiu, Linlu, et al.
Published: (2023)
by: Qiu, Linlu, et al.
Published: (2023)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
by: Han, Xiaochuang, et al.
Published: (2023)
by: Han, Xiaochuang, et al.
Published: (2023)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning
by: Jang, Min, et al.
Published: (2026)
by: Jang, Min, et al.
Published: (2026)
Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia
by: Samir, Farhan, et al.
Published: (2024)
by: Samir, Farhan, et al.
Published: (2024)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization
by: Ahia, Orevaoghene, et al.
Published: (2024)
by: Ahia, Orevaoghene, et al.
Published: (2024)
Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
by: Ahuja, Kabir, et al.
Published: (2026)
by: Ahuja, Kabir, et al.
Published: (2026)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
by: Wu, Addison J., et al.
Published: (2026)
by: Wu, Addison J., et al.
Published: (2026)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
by: Lu, Ximing, et al.
Published: (2024)
by: Lu, Ximing, et al.
Published: (2024)
Are Language Models Sensitive to Morally Irrelevant Distractors?
by: Shaw, Andrew, et al.
Published: (2026)
by: Shaw, Andrew, et al.
Published: (2026)
In-Context Learning through the Bayesian Prism
by: Panwar, Madhur, et al.
Published: (2023)
by: Panwar, Madhur, et al.
Published: (2023)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Don't Throw Away Your Pretrained Model
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
Similar Items
-
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023) -
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
by: Faisal, Fahim, et al.
Published: (2024) -
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025) -
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
by: Sclar, Melanie, et al.
Published: (2024) -
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
by: Ahuja, Kabir, et al.
Published: (2024)