Do LLMs Signal When They're Right? Evidence from Neuron Agreement
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Kang, Wang, Yaoning, Xiong, Kai, Feng, Zhuoka, Sun, Wenhe, Chen, Haotian, Cao, Yixin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Thinking Out Loud: Do Reasoning Models Know When They're Right?
por: Zeng, Qingcheng, et al.
Publicado: (2025)
por: Zeng, Qingcheng, et al.
Publicado: (2025)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
por: Zhang, Anqi, et al.
Publicado: (2025)
por: Zhang, Anqi, et al.
Publicado: (2025)
Do Language Models Know When They're Hallucinating References?
por: Agrawal, Ayush, et al.
Publicado: (2023)
por: Agrawal, Ayush, et al.
Publicado: (2023)
NEX: Neuron Explore-Exploit Scoring for Label-Free Chain-of-Thought Selection and Model Ranking
por: Chen, Kang, et al.
Publicado: (2026)
por: Chen, Kang, et al.
Publicado: (2026)
Do Androids Know They're Only Dreaming of Electric Sheep?
por: CH-Wang, Sky, et al.
Publicado: (2023)
por: CH-Wang, Sky, et al.
Publicado: (2023)
ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging
por: Feng, Zhuoka, et al.
Publicado: (2026)
por: Feng, Zhuoka, et al.
Publicado: (2026)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
por: Cao, Yixin, et al.
Publicado: (2025)
por: Cao, Yixin, et al.
Publicado: (2025)
Thinking Traps in Long Chain-of-Thought: A Measurable Study and Trap-Aware Adaptive Restart
por: Chen, Kang, et al.
Publicado: (2026)
por: Chen, Kang, et al.
Publicado: (2026)
Cogs in a Machine, Doing What They're Meant to Do -- The AMI Submission to the WMT24 General Translation Task
por: Jasonarson, Atli, et al.
Publicado: (2024)
por: Jasonarson, Atli, et al.
Publicado: (2024)
EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization
por: Wang, Yaoning, et al.
Publicado: (2025)
por: Wang, Yaoning, et al.
Publicado: (2025)
Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer
por: Cui, Chenhang, et al.
Publicado: (2026)
por: Cui, Chenhang, et al.
Publicado: (2026)
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
por: Burleigh, Tyler
Publicado: (2026)
por: Burleigh, Tyler
Publicado: (2026)
Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
por: Chen, Kang, et al.
Publicado: (2025)
por: Chen, Kang, et al.
Publicado: (2025)
Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs
por: Pan, Haowen, et al.
Publicado: (2025)
por: Pan, Haowen, et al.
Publicado: (2025)
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
por: Vaddi, Snehit, et al.
Publicado: (2026)
por: Vaddi, Snehit, et al.
Publicado: (2026)
Intuitive or Dependent? Investigating LLMs' Behavior Style to Conflicting Prompts
por: Ying, Jiahao, et al.
Publicado: (2023)
por: Ying, Jiahao, et al.
Publicado: (2023)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
por: Gao, Cheng, et al.
Publicado: (2025)
por: Gao, Cheng, et al.
Publicado: (2025)
Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers
por: Pan, Haowen, et al.
Publicado: (2023)
por: Pan, Haowen, et al.
Publicado: (2023)
The Realignment Problem: When Right becomes Wrong in LLMs
por: Sharma, Aakash Sen, et al.
Publicado: (2025)
por: Sharma, Aakash Sen, et al.
Publicado: (2025)
The Knowledge Microscope: Features as Better Analytical Lenses than Neurons
por: Chen, Yuheng, et al.
Publicado: (2025)
por: Chen, Yuheng, et al.
Publicado: (2025)
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
por: Wang, Haoran, et al.
Publicado: (2026)
por: Wang, Haoran, et al.
Publicado: (2026)
Readability-guided Idiom-aware Sentence Simplification (RISS) for Chinese
por: Zhang, Jingshen, et al.
Publicado: (2024)
por: Zhang, Jingshen, et al.
Publicado: (2024)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
por: Li, Xinze, et al.
Publicado: (2024)
por: Li, Xinze, et al.
Publicado: (2024)
Do LLMs Encode Frame Semantics? Evidence from Frame Identification
por: Chundru, Jayanth Krishna, et al.
Publicado: (2025)
por: Chundru, Jayanth Krishna, et al.
Publicado: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
por: Samway, Keenan, et al.
Publicado: (2026)
por: Samway, Keenan, et al.
Publicado: (2026)
Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
por: Liang, Xun, et al.
Publicado: (2025)
por: Liang, Xun, et al.
Publicado: (2025)
"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
por: Yu, Ye, et al.
Publicado: (2026)
por: Yu, Ye, et al.
Publicado: (2026)
MEMLA: Enhancing Multilingual Knowledge Editing with Neuron-Masked Low-Rank Adaptation
por: Xie, Jiakuan, et al.
Publicado: (2024)
por: Xie, Jiakuan, et al.
Publicado: (2024)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
por: Zhang, Zhuoxuan, et al.
Publicado: (2025)
por: Zhang, Zhuoxuan, et al.
Publicado: (2025)
Diagnosing and Remedying Knowledge Deficiencies in LLMs via Label-free Curricular Meaningful Learning
por: Xiong, Kai, et al.
Publicado: (2024)
por: Xiong, Kai, et al.
Publicado: (2024)
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models
por: Yu, Ye, et al.
Publicado: (2025)
por: Yu, Ye, et al.
Publicado: (2025)
One Mind, Many Tongues: A Deep Dive into Language-Agnostic Knowledge Neurons in Large Language Models
por: Cao, Pengfei, et al.
Publicado: (2024)
por: Cao, Pengfei, et al.
Publicado: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
por: Kidder, William, et al.
Publicado: (2024)
por: Kidder, William, et al.
Publicado: (2024)
LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
por: Ahtisham, Bakhtawar, et al.
Publicado: (2026)
por: Ahtisham, Bakhtawar, et al.
Publicado: (2026)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
por: Ni, Shiyu, et al.
Publicado: (2024)
por: Ni, Shiyu, et al.
Publicado: (2024)
Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language Models
por: Chen, Yuheng, et al.
Publicado: (2024)
por: Chen, Yuheng, et al.
Publicado: (2024)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
por: Wan, Yixin, et al.
Publicado: (2024)
por: Wan, Yixin, et al.
Publicado: (2024)
Do the Right Thing, Just Debias! Multi-Category Bias Mitigation Using LLMs
por: Roy, Amartya, et al.
Publicado: (2024)
por: Roy, Amartya, et al.
Publicado: (2024)
LANDeRMT: Detecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation
por: Zhu, Shaolin, et al.
Publicado: (2024)
por: Zhu, Shaolin, et al.
Publicado: (2024)
Ejemplares similares
-
Thinking Out Loud: Do Reasoning Models Know When They're Right?
por: Zeng, Qingcheng, et al.
Publicado: (2025) -
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
por: Zhang, Anqi, et al.
Publicado: (2025) -
Do Language Models Know When They're Hallucinating References?
por: Agrawal, Ayush, et al.
Publicado: (2023) -
NEX: Neuron Explore-Exploit Scoring for Label-Free Chain-of-Thought Selection and Model Ranking
por: Chen, Kang, et al.
Publicado: (2026) -
Do Androids Know They're Only Dreaming of Electric Sheep?
por: CH-Wang, Sky, et al.
Publicado: (2023)