On the Universal Truthfulness Hyperplane Inside LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Junteng, Chen, Shiqi, Cheng, Yu, He, Junxian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
von: Zhou, Ruochen, et al.
Veröffentlicht: (2025)
von: Zhou, Ruochen, et al.
Veröffentlicht: (2025)
In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
von: Chen, Shiqi, et al.
Veröffentlicht: (2024)
von: Chen, Shiqi, et al.
Veröffentlicht: (2024)
On the Perception Bottleneck of VLMs for Chart Understanding
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
Truth is Universal: Robust Detection of Lies in LLMs
von: Bürger, Lennart, et al.
Veröffentlicht: (2024)
von: Bürger, Lennart, et al.
Veröffentlicht: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
Training-free Truthfulness Detection via Value Vectors in LLMs
von: Liu, Runheng, et al.
Veröffentlicht: (2025)
von: Liu, Runheng, et al.
Veröffentlicht: (2025)
Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
von: Jin, Xiangqi, et al.
Veröffentlicht: (2025)
von: Jin, Xiangqi, et al.
Veröffentlicht: (2025)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
von: Tong, Zhao, et al.
Veröffentlicht: (2026)
von: Tong, Zhao, et al.
Veröffentlicht: (2026)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
Testing the Limits of Truth Directions in LLMs
von: Poulis, Angelos, et al.
Veröffentlicht: (2026)
von: Poulis, Angelos, et al.
Veröffentlicht: (2026)
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
von: Bao, Yuntai, et al.
Veröffentlicht: (2025)
von: Bao, Yuntai, et al.
Veröffentlicht: (2025)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
von: Adarsh, Shivam, et al.
Veröffentlicht: (2026)
von: Adarsh, Shivam, et al.
Veröffentlicht: (2026)
Reasoning Gets Harder for LLMs Inside A Dialogue
von: Kartáč, Ivan, et al.
Veröffentlicht: (2026)
von: Kartáč, Ivan, et al.
Veröffentlicht: (2026)
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
Truth Neurons
von: Li, Haohang, et al.
Veröffentlicht: (2025)
von: Li, Haohang, et al.
Veröffentlicht: (2025)
Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents
von: He, Haorui, et al.
Veröffentlicht: (2025)
von: He, Haorui, et al.
Veröffentlicht: (2025)
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
von: Karia, Rushang, et al.
Veröffentlicht: (2024)
von: Karia, Rushang, et al.
Veröffentlicht: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2023)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2023)
Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
Debating with More Persuasive LLMs Leads to More Truthful Answers
von: Khan, Akbir, et al.
Veröffentlicht: (2024)
von: Khan, Akbir, et al.
Veröffentlicht: (2024)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
von: Chatrath, Veronica, et al.
Veröffentlicht: (2024)
von: Chatrath, Veronica, et al.
Veröffentlicht: (2024)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
von: Agrawal, Tanmay
Veröffentlicht: (2025)
von: Agrawal, Tanmay
Veröffentlicht: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
Ground Truth Generation for Multilingual Historical NLP using LLMs
von: Gladstone, Clovis, et al.
Veröffentlicht: (2025)
von: Gladstone, Clovis, et al.
Veröffentlicht: (2025)
Diving into Self-Evolving Training for Multimodal Reasoning
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
High-Dimensional Interlingual Representations of Large Language Models
von: Wilie, Bryan, et al.
Veröffentlicht: (2025)
von: Wilie, Bryan, et al.
Veröffentlicht: (2025)
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
von: Bini, Massimo, et al.
Veröffentlicht: (2024)
von: Bini, Massimo, et al.
Veröffentlicht: (2024)
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs
von: Yaldiz, Duygu Nur, et al.
Veröffentlicht: (2025)
von: Yaldiz, Duygu Nur, et al.
Veröffentlicht: (2025)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
von: Matys, Piotr, et al.
Veröffentlicht: (2025)
von: Matys, Piotr, et al.
Veröffentlicht: (2025)
Internalizing World Models via Self-Play Finetuning for Agentic RL
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
von: Nguyen, Thi-Nhung, et al.
Veröffentlicht: (2026)
von: Nguyen, Thi-Nhung, et al.
Veröffentlicht: (2026)
Generalist Reward Models: Found Inside Large Language Models
von: Li, Yi-Chen, et al.
Veröffentlicht: (2025)
von: Li, Yi-Chen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
von: Zhou, Ruochen, et al.
Veröffentlicht: (2025) -
In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
von: Chen, Shiqi, et al.
Veröffentlicht: (2024) -
On the Perception Bottleneck of VLMs for Chart Understanding
von: Liu, Junteng, et al.
Veröffentlicht: (2025) -
Truth is Universal: Robust Detection of Lies in LLMs
von: Bürger, Lennart, et al.
Veröffentlicht: (2024) -
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)