Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Bao, Yuntai, Zhang, Xuhong, Du, Tianyu, Zhao, Xinkui, Feng, Zhengwen, Peng, Hao, Yin, Jianwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scalable Multi-Stage Influence Function for Large Language Models via Eigenvalue-Corrected Kronecker-Factored Parameterization
por: Bao, Yuntai, et al.
Publicado: (2025)
por: Bao, Yuntai, et al.
Publicado: (2025)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
por: Zhang, Yucheng, et al.
Publicado: (2024)
por: Zhang, Yucheng, et al.
Publicado: (2024)
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025)
por: Azizian, Waiss, et al.
Publicado: (2025)
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
por: Bao, Yuntai, et al.
Publicado: (2026)
por: Bao, Yuntai, et al.
Publicado: (2026)
Testing the Limits of Truth Directions in LLMs
por: Poulis, Angelos, et al.
Publicado: (2026)
por: Poulis, Angelos, et al.
Publicado: (2026)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
por: Fu, Yao, et al.
Publicado: (2025)
por: Fu, Yao, et al.
Publicado: (2025)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
por: Adarsh, Shivam, et al.
Publicado: (2026)
por: Adarsh, Shivam, et al.
Publicado: (2026)
Logical Form and Truth-Conditions
por: Andrea IACONA
Publicado: (2013)
por: Andrea IACONA
Publicado: (2013)
Truth Claims Across Media
Publicado: (2024)
Publicado: (2024)
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
por: Karia, Rushang, et al.
Publicado: (2024)
por: Karia, Rushang, et al.
Publicado: (2024)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
Safeguarding the Truth of High-Value Price Oracle Task: A Dynamically Adjusted Truth Discovery Method
por: Xian, Youquan, et al.
Publicado: (2024)
por: Xian, Youquan, et al.
Publicado: (2024)
The Muses of Truth and Transformation
por: Chinen, Allan B.
Publicado: (2024)
por: Chinen, Allan B.
Publicado: (2024)
Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages
por: Bayes, Edward, et al.
Publicado: (2024)
por: Bayes, Edward, et al.
Publicado: (2024)
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
por: Bao, Yuntai, et al.
Publicado: (2026)
por: Bao, Yuntai, et al.
Publicado: (2026)
ERA-CoT: Improving Chain-of-Thought through Entity Relationship Analysis
por: Liu, Yanming, et al.
Publicado: (2024)
por: Liu, Yanming, et al.
Publicado: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
por: Jaiswal, Ajay, et al.
Publicado: (2023)
por: Jaiswal, Ajay, et al.
Publicado: (2023)
CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment
por: Li, Qinfeng, et al.
Publicado: (2024)
por: Li, Qinfeng, et al.
Publicado: (2024)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
por: Kumar, Aayush
Publicado: (2025)
por: Kumar, Aayush
Publicado: (2025)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
por: Li, Qinfeng, et al.
Publicado: (2024)
por: Li, Qinfeng, et al.
Publicado: (2024)
Tool-Planner: Task Planning with Clusters across Multiple Tools
por: Liu, Yanming, et al.
Publicado: (2024)
por: Liu, Yanming, et al.
Publicado: (2024)
On the Universal Truthfulness Hyperplane Inside LLMs
por: Liu, Junteng, et al.
Publicado: (2024)
por: Liu, Junteng, et al.
Publicado: (2024)
Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D Generation
por: Li, Ziying, et al.
Publicado: (2025)
por: Li, Ziying, et al.
Publicado: (2025)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
por: Wang, Hanyu, et al.
Publicado: (2025)
por: Wang, Hanyu, et al.
Publicado: (2025)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
por: Zhang, Shaolei, et al.
Publicado: (2024)
por: Zhang, Shaolei, et al.
Publicado: (2024)
Love the Truth, the Whole Truth, and the Truth about Everything. An Interview with Josef Seifert
por: Rodrigo Guerra López
Publicado: (2014)
por: Rodrigo Guerra López
Publicado: (2014)
MIDAS: Modeling Ground-Truth Distributions with Dark Knowledge for Domain Generalized Stereo Matching
por: Xu, Peng, et al.
Publicado: (2025)
por: Xu, Peng, et al.
Publicado: (2025)
Queries With Exact Truth Values in Paraconsistent Description Logics
por: Bienvenu, Meghyn, et al.
Publicado: (2024)
por: Bienvenu, Meghyn, et al.
Publicado: (2024)
Accurate Table Question Answering with Accessible LLMs
por: Jiang, Yangfan, et al.
Publicado: (2026)
por: Jiang, Yangfan, et al.
Publicado: (2026)
Truth
por: Cubitt, Sean
Publicado: (2024)
por: Cubitt, Sean
Publicado: (2024)
Ground Truth Generation for Multilingual Historical NLP using LLMs
por: Gladstone, Clovis, et al.
Publicado: (2025)
por: Gladstone, Clovis, et al.
Publicado: (2025)
The Truth, the Whole Truth, and Nothing but the Truth: Automatic Visualization Evaluation from Reconstruction Quality
por: Bujack, Roxana, et al.
Publicado: (2026)
por: Bujack, Roxana, et al.
Publicado: (2026)
Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
Compositional Consistency-Guided Decoding for Three-Way Logical Question Answering
por: Huang, Tianyi, et al.
Publicado: (2026)
por: Huang, Tianyi, et al.
Publicado: (2026)
Truthful Aggregation of LLMs with an Application to Online Advertising
por: Soumalias, Ermis, et al.
Publicado: (2024)
por: Soumalias, Ermis, et al.
Publicado: (2024)
Truth is Universal: Robust Detection of Lies in LLMs
por: Bürger, Lennart, et al.
Publicado: (2024)
por: Bürger, Lennart, et al.
Publicado: (2024)
RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-Feedback
por: Liu, Yanming, et al.
Publicado: (2024)
por: Liu, Yanming, et al.
Publicado: (2024)
Neural Quantum States in Variational Monte Carlo Method: A Brief Summary
por: Song, Yuntai
Publicado: (2024)
por: Song, Yuntai
Publicado: (2024)
Truth Knows No Language: Evaluating Truthfulness Beyond English
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
por: Ameen, Fathima, et al.
Publicado: (2026)
por: Ameen, Fathima, et al.
Publicado: (2026)
Ejemplares similares
-
Scalable Multi-Stage Influence Function for Large Language Models via Eigenvalue-Corrected Kronecker-Factored Parameterization
por: Bao, Yuntai, et al.
Publicado: (2025) -
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
por: Zhang, Yucheng, et al.
Publicado: (2024) -
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025) -
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
por: Bao, Yuntai, et al.
Publicado: (2026) -
Testing the Limits of Truth Directions in LLMs
por: Poulis, Angelos, et al.
Publicado: (2026)