Cutting Off the Head Ends the Conflict: A Mechanism for Interpreting and Mitigating Knowledge Conflicts in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Zhuoran, Cao, Pengfei, Yuan, Hongbang, Chen, Yubo, Xu, Jiexin, Li, Huaijun, Jiang, Xiaojian, Liu, Kang, Zhao, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tug-of-War Between Knowledge: Exploring and Resolving Knowledge Conflicts in Retrieval-Augmented Language Models
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
by: He, Zhitao, et al.
Published: (2024)
by: He, Zhitao, et al.
Published: (2024)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling
by: Li, Jiachun, et al.
Published: (2025)
by: Li, Jiachun, et al.
Published: (2025)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
LINKED: Eliciting, Filtering and Integrating Knowledge in Large Language Model for Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts
by: Tao, Zhen, et al.
Published: (2026)
by: Tao, Zhen, et al.
Published: (2026)
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
by: Pham, Quang Hieu, et al.
Published: (2024)
by: Pham, Quang Hieu, et al.
Published: (2024)
Knowledge Conflicts for LLMs: A Survey
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
by: Men, Tianyi, et al.
Published: (2024)
by: Men, Tianyi, et al.
Published: (2024)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
by: Jin, Zhuoran, et al.
Published: (2025)
by: Jin, Zhuoran, et al.
Published: (2025)
One Mind, Many Tongues: A Deep Dive into Language-Agnostic Knowledge Neurons in Large Language Models
by: Cao, Pengfei, et al.
Published: (2024)
by: Cao, Pengfei, et al.
Published: (2024)
SA-CAISR: Stage-Adaptive and Conflict-Aware Incremental Sequential Recommendation
by: Song, Xiaomeng, et al.
Published: (2026)
by: Song, Xiaomeng, et al.
Published: (2026)
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
FRESCO: Benchmarking and Optimizing Re-rankers for Evolving Semantic Conflict in Retrieval-Augmented Generation
by: An, Sohyun, et al.
Published: (2026)
by: An, Sohyun, et al.
Published: (2026)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
ArbGraph: Conflict-Aware Evidence Arbitration for Reliable Long-Form Retrieval-Augmented Generation
by: Niu, Qingying, et al.
Published: (2026)
by: Niu, Qingying, et al.
Published: (2026)
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
IF-GEO: Conflict-Aware Instruction Fusion for Multi-Query Generative Engine Optimization
by: Zhou, Heyang, et al.
Published: (2026)
by: Zhou, Heyang, et al.
Published: (2026)
Temporal Fact Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN
by: Dey, Ritajit, et al.
Published: (2026)
by: Dey, Ritajit, et al.
Published: (2026)
Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation
by: Zhang, Jiang, et al.
Published: (2025)
by: Zhang, Jiang, et al.
Published: (2025)
Conflicts of Interest in Published NLP Research 2000-2024
by: Bosten, Maarten, et al.
Published: (2025)
by: Bosten, Maarten, et al.
Published: (2025)
Digital Gatekeeping: An Audit of Search Engine Results shows tailoring of queries on the Israel-Palestine Conflict
by: Damião, Íris, et al.
Published: (2025)
by: Damião, Íris, et al.
Published: (2025)
Effective Knowledge Transfer for Multi-Task Recommendation Models
by: Cai, Guohao, et al.
Published: (2026)
by: Cai, Guohao, et al.
Published: (2026)
FactCHD: Benchmarking Fact-Conflicting Hallucination Detection
by: Chen, Xiang, et al.
Published: (2023)
by: Chen, Xiang, et al.
Published: (2023)
Towards Mitigating Dimensional Collapse of Representations in Collaborative Filtering
by: Chen, Huiyuan, et al.
Published: (2023)
by: Chen, Huiyuan, et al.
Published: (2023)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Aligning Large Language Models with Recommendation Knowledge
by: Cao, Yuwei, et al.
Published: (2024)
by: Cao, Yuwei, et al.
Published: (2024)
On Mitigating Data Sparsity in Conversational Recommender Systems
by: Zhang, Sixiao, et al.
Published: (2025)
by: Zhang, Sixiao, et al.
Published: (2025)
Interpret and Control Dense Retrieval with Sparse Latent Features
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
An Analysis on Matching Mechanisms and Token Pruning for Late-interaction Models
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
From Head to Tail: Asymmetric Knowledge Transfer in Long-tail Recommendation with Generative Semantic IDs
by: Yan, Chenyi, et al.
Published: (2026)
by: Yan, Chenyi, et al.
Published: (2026)
ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor
by: Pan, Taiheng
Published: (2026)
by: Pan, Taiheng
Published: (2026)
GDLLM: A Global Distance-aware Modeling Approach Based on Large Language Models for Event Temporal Relation Extraction
by: Zhao, Jie, et al.
Published: (2025)
by: Zhao, Jie, et al.
Published: (2025)
When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
Similar Items
-
Tug-of-War Between Knowledge: Exploring and Resolving Knowledge Conflicts in Retrieval-Augmented Language Models
by: Jin, Zhuoran, et al.
Published: (2024) -
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024) -
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024) -
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
by: He, Zhitao, et al.
Published: (2024) -
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)