The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tong, Zekai, Xu, Ruiyao, Shrivastava, Aryan, Tan, Chenhao, Holtzman, Ari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
Prompting as Scientific Inquiry
von: Holtzman, Ari, et al.
Veröffentlicht: (2025)
von: Holtzman, Ari, et al.
Veröffentlicht: (2025)
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
AbsenceBench: Language Models Can't Tell What's Missing
von: Fu, Harvey Yiyun, et al.
Veröffentlicht: (2025)
von: Fu, Harvey Yiyun, et al.
Veröffentlicht: (2025)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
Iterative Finetuning is Mostly Idempotent
von: Roe, Zephaniah, et al.
Veröffentlicht: (2026)
von: Roe, Zephaniah, et al.
Veröffentlicht: (2026)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
von: Li, Margaret, et al.
Veröffentlicht: (2024)
von: Li, Margaret, et al.
Veröffentlicht: (2024)
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
von: Cho, Yong-eun
Veröffentlicht: (2026)
von: Cho, Yong-eun
Veröffentlicht: (2026)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Resource-Aware Arabic LLM Creation: Model Adaptation, Integration, and Multi-Domain Testing
von: Aryan, Prakash
Veröffentlicht: (2024)
von: Aryan, Prakash
Veröffentlicht: (2024)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
von: Wu, Juntong, et al.
Veröffentlicht: (2025)
von: Wu, Juntong, et al.
Veröffentlicht: (2025)
AgentStealth: Reinforcing Large Language Model for Anonymizing User-generated Text
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
von: Keluskar, Aryan, et al.
Veröffentlicht: (2024)
von: Keluskar, Aryan, et al.
Veröffentlicht: (2024)
Context Discipline and Performance Correlation: Analyzing LLM Performance and Quality Degradation Under Varying Context Lengths
von: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Veröffentlicht: (2025)
von: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Veröffentlicht: (2025)
AI as Entertainment
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
Cat-DPO: Category-Adaptive Safety Alignment
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length
von: Chen, Jingxuan, et al.
Veröffentlicht: (2026)
von: Chen, Jingxuan, et al.
Veröffentlicht: (2026)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
Training With "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-context Tasks
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
von: Dai, Sunhao, et al.
Veröffentlicht: (2024)
von: Dai, Sunhao, et al.
Veröffentlicht: (2024)
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
von: Ahmed, Zeeshan, et al.
Veröffentlicht: (2025)
von: Ahmed, Zeeshan, et al.
Veröffentlicht: (2025)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery
von: Gu, Hongchao, et al.
Veröffentlicht: (2025)
von: Gu, Hongchao, et al.
Veröffentlicht: (2025)
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
von: Ma, Guangyuan, et al.
Veröffentlicht: (2025)
von: Ma, Guangyuan, et al.
Veröffentlicht: (2025)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
von: Zhou, Karen, et al.
Veröffentlicht: (2025)
von: Zhou, Karen, et al.
Veröffentlicht: (2025)
MUSE: Machine Unlearning Six-Way Evaluation for Language Models
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry
von: Liao, Hsien-Jyh
Veröffentlicht: (2026)
von: Liao, Hsien-Jyh
Veröffentlicht: (2026)
Mast Kalandar at SemEval-2024 Task 8: On the Trail of Textual Origins: RoBERTa-BiLSTM Approach to Detect AI-Generated Text
von: Bafna, Jainit Sushil, et al.
Veröffentlicht: (2024)
von: Bafna, Jainit Sushil, et al.
Veröffentlicht: (2024)
VBART: The Turkish LLM
von: Turker, Meliksah, et al.
Veröffentlicht: (2024)
von: Turker, Meliksah, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025) -
Know Thyself? On the Incapability and Implications of AI Self-Recognition
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025) -
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026) -
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026) -
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)