Testing the Limits of Truth Directions in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Poulis, Angelos, Crovella, Mark, Terzi, Evimaria |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Transformer-based Language Models for Reasoning in the Description Logic ALCQ
por: Poulis, Angelos, et al.
Publicado: (2024)
por: Poulis, Angelos, et al.
Publicado: (2024)
Transformers in the Service of Description Logic-based Contexts
por: Poulis, Angelos, et al.
Publicado: (2023)
por: Poulis, Angelos, et al.
Publicado: (2023)
Sparse Attention Decomposition Applied to Circuit Tracing
por: Franco, Gabriel, et al.
Publicado: (2024)
por: Franco, Gabriel, et al.
Publicado: (2024)
Team Formation amidst Conflicts
por: Nikolaou, Iasonas, et al.
Publicado: (2024)
por: Nikolaou, Iasonas, et al.
Publicado: (2024)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
por: Adarsh, Shivam, et al.
Publicado: (2026)
por: Adarsh, Shivam, et al.
Publicado: (2026)
Truth is Universal: Robust Detection of Lies in LLMs
por: Bürger, Lennart, et al.
Publicado: (2024)
por: Bürger, Lennart, et al.
Publicado: (2024)
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
por: Karia, Rushang, et al.
Publicado: (2024)
por: Karia, Rushang, et al.
Publicado: (2024)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
por: Agrawal, Tanmay
Publicado: (2025)
por: Agrawal, Tanmay
Publicado: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
por: Barkett, Emilio, et al.
Publicado: (2025)
por: Barkett, Emilio, et al.
Publicado: (2025)
Debating with More Persuasive LLMs Leads to More Truthful Answers
por: Khan, Akbir, et al.
Publicado: (2024)
por: Khan, Akbir, et al.
Publicado: (2024)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
por: Chatrath, Veronica, et al.
Publicado: (2024)
por: Chatrath, Veronica, et al.
Publicado: (2024)
Ground Truth Generation for Multilingual Historical NLP using LLMs
por: Gladstone, Clovis, et al.
Publicado: (2025)
por: Gladstone, Clovis, et al.
Publicado: (2025)
Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs
por: Servedio, Giovanni, et al.
Publicado: (2025)
por: Servedio, Giovanni, et al.
Publicado: (2025)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
por: Matys, Piotr, et al.
Publicado: (2025)
por: Matys, Piotr, et al.
Publicado: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
por: Ameen, Fathima, et al.
Publicado: (2026)
por: Ameen, Fathima, et al.
Publicado: (2026)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
por: Khatun, Aisha, et al.
Publicado: (2024)
por: Khatun, Aisha, et al.
Publicado: (2024)
Truth Knows No Language: Evaluating Truthfulness Beyond English
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
Multi-Agent LLMs for Generating Research Limitations
por: Azher, Ibrahim Al, et al.
Publicado: (2025)
por: Azher, Ibrahim Al, et al.
Publicado: (2025)
FACEGroup: Feasible and Actionable Counterfactual Explanations for Group Fairness
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
por: Chen, Zhongzhi, et al.
Publicado: (2023)
por: Chen, Zhongzhi, et al.
Publicado: (2023)
To Tell The Truth: Language of Deception and Language Models
por: Hazra, Sanchaita, et al.
Publicado: (2023)
por: Hazra, Sanchaita, et al.
Publicado: (2023)
Recon, Answer, Verify: Agents in Search of Truth
por: Shukla, Satyam, et al.
Publicado: (2025)
por: Shukla, Satyam, et al.
Publicado: (2025)
Truth, Trust, and Trouble: Medical AI on the Edge
por: Azeez, Mohammad Anas, et al.
Publicado: (2025)
por: Azeez, Mohammad Anas, et al.
Publicado: (2025)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
por: Luo, Wen, et al.
Publicado: (2026)
por: Luo, Wen, et al.
Publicado: (2026)
Improving Fairness in LLMs Through Testing-Time Adversaries
por: Gregio, Isabela Pereira, et al.
Publicado: (2025)
por: Gregio, Isabela Pereira, et al.
Publicado: (2025)
On the Relationship between Truth and Political Bias in Language Models
por: Fulay, Suyash, et al.
Publicado: (2024)
por: Fulay, Suyash, et al.
Publicado: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
por: Du, Hongzhe, et al.
Publicado: (2025)
por: Du, Hongzhe, et al.
Publicado: (2025)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
por: Zhang, Xiang, et al.
Publicado: (2025)
por: Zhang, Xiang, et al.
Publicado: (2025)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
por: Wang, Hanyu, et al.
Publicado: (2025)
por: Wang, Hanyu, et al.
Publicado: (2025)
Directional Optimization Asymmetry in Transformers: A Synthetic Stress Test
por: Sahasrabudhe, Mihir
Publicado: (2025)
por: Sahasrabudhe, Mihir
Publicado: (2025)
Mind Reading or Misreading? LLMs on the Big Five Personality Test
por: Di Cursi, Francesco, et al.
Publicado: (2025)
por: Di Cursi, Francesco, et al.
Publicado: (2025)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
por: Wu, Tianyi, et al.
Publicado: (2025)
por: Wu, Tianyi, et al.
Publicado: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
por: Guo, Ruiling, et al.
Publicado: (2025)
por: Guo, Ruiling, et al.
Publicado: (2025)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
por: Zhang, Shaolei, et al.
Publicado: (2024)
por: Zhang, Shaolei, et al.
Publicado: (2024)
From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
por: Tang, Luoxi, et al.
Publicado: (2025)
por: Tang, Luoxi, et al.
Publicado: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
por: Nadeem, Afrozah, et al.
Publicado: (2025)
por: Nadeem, Afrozah, et al.
Publicado: (2025)
Ejemplares similares
-
Transformer-based Language Models for Reasoning in the Description Logic ALCQ
por: Poulis, Angelos, et al.
Publicado: (2024) -
Transformers in the Service of Description Logic-based Contexts
por: Poulis, Angelos, et al.
Publicado: (2023) -
Sparse Attention Decomposition Applied to Circuit Tracing
por: Franco, Gabriel, et al.
Publicado: (2024) -
Team Formation amidst Conflicts
por: Nikolaou, Iasonas, et al.
Publicado: (2024) -
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
por: Adarsh, Shivam, et al.
Publicado: (2026)