AbsenceBench: Language Models Can't Tell What's Missing
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Harvey Yiyun, Shrivastava, Aryan, Moore, Jared, West, Peter, Tan, Chenhao, Holtzman, Ari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
by: Bai, Xiaoyan, et al.
Published: (2025)
by: Bai, Xiaoyan, et al.
Published: (2025)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
by: Tong, Zekai, et al.
Published: (2026)
by: Tong, Zekai, et al.
Published: (2026)
Prompting as Scientific Inquiry
by: Holtzman, Ari, et al.
Published: (2025)
by: Holtzman, Ari, et al.
Published: (2025)
Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
by: Holtzman, Ari, et al.
Published: (2026)
by: Holtzman, Ari, et al.
Published: (2026)
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
by: Li, Margaret, et al.
Published: (2024)
by: Li, Margaret, et al.
Published: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
Benchmarks as Microscopes: A Call for Model Metrology
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
by: Shrivastava, Aryan, et al.
Published: (2024)
by: Shrivastava, Aryan, et al.
Published: (2024)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
by: Liu, Yongkang, et al.
Published: (2023)
by: Liu, Yongkang, et al.
Published: (2023)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
by: Gligoric, Kristina, et al.
Published: (2024)
by: Gligoric, Kristina, et al.
Published: (2024)
Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
by: Sui, Peiqi, et al.
Published: (2026)
by: Sui, Peiqi, et al.
Published: (2026)
What Artificial Neural Networks Can Tell Us About Human Language Acquisition
by: Warstadt, Alex, et al.
Published: (2022)
by: Warstadt, Alex, et al.
Published: (2022)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
by: He, Jiahang, et al.
Published: (2025)
by: He, Jiahang, et al.
Published: (2025)
What is Wrong with Language Models that Can Not Tell a Story?
by: Yamshchikov, Ivan P., et al.
Published: (2022)
by: Yamshchikov, Ivan P., et al.
Published: (2022)
Subliminal Learning is a LoRA Artifact
by: Nief, Todd, et al.
Published: (2026)
by: Nief, Todd, et al.
Published: (2026)
You Can't Fight in Here! This is BBS!
by: Futrell, Richard, et al.
Published: (2026)
by: Futrell, Richard, et al.
Published: (2026)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Semantic Deception: When Reasoning Models Can't Compute an Addition
by: de Leeuw, Nathaniël, et al.
Published: (2025)
by: de Leeuw, Nathaniël, et al.
Published: (2025)
What Can String Probability Tell Us About Grammaticality?
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
by: Ji, Xu, et al.
Published: (2024)
by: Ji, Xu, et al.
Published: (2024)
Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
Can Language Models Recognize Convincing Arguments?
by: Rescala, Paula, et al.
Published: (2024)
by: Rescala, Paula, et al.
Published: (2024)
Can Constructions "SCAN" Compositionality ?
by: Katrapati, Ganesh, et al.
Published: (2025)
by: Katrapati, Ganesh, et al.
Published: (2025)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
by: Baldelli, Davide, et al.
Published: (2026)
by: Baldelli, Davide, et al.
Published: (2026)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
Are Large Language Models Consistent over Value-laden Questions?
by: Moore, Jared, et al.
Published: (2024)
by: Moore, Jared, et al.
Published: (2024)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
by: Ulmer, Dennis, et al.
Published: (2025)
by: Ulmer, Dennis, et al.
Published: (2025)
What Is Missing: Interpretable Ratings for Large Language Model Outputs
by: Stranges, Nicholas, et al.
Published: (2026)
by: Stranges, Nicholas, et al.
Published: (2026)
Forking Paths in Neural Text Generation
by: Bigelow, Eric, et al.
Published: (2024)
by: Bigelow, Eric, et al.
Published: (2024)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)
by: Riachi, Roland, et al.
Published: (2025)
SysBench: Can Large Language Models Follow System Messages?
by: Qin, Yanzhao, et al.
Published: (2024)
by: Qin, Yanzhao, et al.
Published: (2024)
Similar Items
-
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026) -
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025) -
Know Thyself? On the Incapability and Implications of AI Self-Recognition
by: Bai, Xiaoyan, et al.
Published: (2025) -
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
by: Tong, Zekai, et al.
Published: (2026) -
Prompting as Scientific Inquiry
by: Holtzman, Ari, et al.
Published: (2025)