Biases in the Blind Spot: Detecting What LLMs Fail to Mention
Fuente:
arXiv
Saved in:
| Main Authors: | Arcuschin, Iván, Chanin, David, Garriga-Alonso, Adrià, Camburu, Oana-Maria |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026)
by: Chanin, David, et al.
Published: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026)
by: Wang, Atticus, et al.
Published: (2026)
Identifying Linear Relational Concepts in Large Language Models
by: Chanin, David, et al.
Published: (2023)
by: Chanin, David, et al.
Published: (2023)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
by: Gupta, Rohan, et al.
Published: (2024)
by: Gupta, Rohan, et al.
Published: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
by: Chanin, David
Published: (2026)
by: Chanin, David
Published: (2026)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
by: Taufeeque, Mohammad, et al.
Published: (2025)
by: Taufeeque, Mohammad, et al.
Published: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
by: Bush, Thomas, et al.
Published: (2025)
by: Bush, Thomas, et al.
Published: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
DiFR: Inference Verification Despite Nondeterminism
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Planning in a recurrent neural network that plays Sokoban
by: Taufeeque, Mohammad, et al.
Published: (2024)
by: Taufeeque, Mohammad, et al.
Published: (2024)
Inference-Time Toxicity Mitigation in Protein Language Models
by: Burda, Manuel Fernández, et al.
Published: (2026)
by: Burda, Manuel Fernández, et al.
Published: (2026)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
AD4RL: Autonomous Driving Benchmarks for Offline Reinforcement Learning with Value-based Dataset
by: Lee, Dongsu, et al.
Published: (2024)
by: Lee, Dongsu, et al.
Published: (2024)
Linguistic Blind Spots of Large Language Models
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
by: Meek, Austin, et al.
Published: (2025)
by: Meek, Austin, et al.
Published: (2025)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
by: Wang, Yongxin, et al.
Published: (2025)
by: Wang, Yongxin, et al.
Published: (2025)
JEL: A Novel Model Linking Knowledge Graph entities to News Mentions
by: Kishelev, Michael, et al.
Published: (2025)
by: Kishelev, Michael, et al.
Published: (2025)
Analyzing the Generalization and Reliability of Steering Vectors
by: Tan, Daniel, et al.
Published: (2024)
by: Tan, Daniel, et al.
Published: (2024)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
Belief Aided Navigation using Bayesian Reinforcement Learning for Avoiding Humans in Blind Spots
by: Kim, Jinyeob, et al.
Published: (2024)
by: Kim, Jinyeob, et al.
Published: (2024)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
by: Ha, SeungBum, et al.
Published: (2025)
by: Ha, SeungBum, et al.
Published: (2025)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Investigating the Indirect Object Identification circuit in Mamba
by: Ensign, Danielle, et al.
Published: (2024)
by: Ensign, Danielle, et al.
Published: (2024)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
by: Tsui, Ken
Published: (2025)
by: Tsui, Ken
Published: (2025)
Message-Passing GNNs Fail to Approximate Sparse Triangular Factorizations
by: Trifonov, Vladislav, et al.
Published: (2025)
by: Trifonov, Vladislav, et al.
Published: (2025)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
by: Yang, Xikang, et al.
Published: (2025)
by: Yang, Xikang, et al.
Published: (2025)
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024)
by: Sun, Jimin, et al.
Published: (2024)
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
by: Wu, David X., et al.
Published: (2025)
by: Wu, David X., et al.
Published: (2025)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024)
by: Öncel, Fırat, et al.
Published: (2024)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
by: Cox, Kyle, et al.
Published: (2026)
by: Cox, Kyle, et al.
Published: (2026)
BSG4Bot: Efficient Bot Detection based on Biased Heterogeneous Subgraphs
by: Miao, Hao, et al.
Published: (2024)
by: Miao, Hao, et al.
Published: (2024)
Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference
by: Yu, Jianxing, et al.
Published: (2024)
by: Yu, Jianxing, et al.
Published: (2024)
Similar Items
-
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026) -
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025) -
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025) -
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026)