Training-free Truthfulness Detection via Value Vectors in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Runheng, Huang, Heyan, Xiao, Xingchen, Wu, Zhijing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
por: Liu, Runheng, et al.
Publicado: (2026)
por: Liu, Runheng, et al.
Publicado: (2026)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
por: Liu, Runheng, et al.
Publicado: (2024)
por: Liu, Runheng, et al.
Publicado: (2024)
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
por: Xiao, Xingchen, et al.
Publicado: (2026)
por: Xiao, Xingchen, et al.
Publicado: (2026)
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
por: Zhou, Youchao, et al.
Publicado: (2024)
por: Zhou, Youchao, et al.
Publicado: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
por: Zhang, Guo-Biao, et al.
Publicado: (2026)
por: Zhang, Guo-Biao, et al.
Publicado: (2026)
EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models
por: Xie, Jincheng, et al.
Publicado: (2026)
por: Xie, Jincheng, et al.
Publicado: (2026)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
por: Lu, Yi-Fan, et al.
Publicado: (2024)
por: Lu, Yi-Fan, et al.
Publicado: (2024)
Truth is Universal: Robust Detection of Lies in LLMs
por: Bürger, Lennart, et al.
Publicado: (2024)
por: Bürger, Lennart, et al.
Publicado: (2024)
On the Universal Truthfulness Hyperplane Inside LLMs
por: Liu, Junteng, et al.
Publicado: (2024)
por: Liu, Junteng, et al.
Publicado: (2024)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
por: Liu, Yuhang, et al.
Publicado: (2026)
por: Liu, Yuhang, et al.
Publicado: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
por: Yao, Jiashu, et al.
Publicado: (2026)
por: Yao, Jiashu, et al.
Publicado: (2026)
ValueSim: Generating Backstories to Model Individual Value Systems
por: Du, Bangde, et al.
Publicado: (2025)
por: Du, Bangde, et al.
Publicado: (2025)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
por: Qiu, Yuli, et al.
Publicado: (2024)
por: Qiu, Yuli, et al.
Publicado: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
por: Yang, Gao, et al.
Publicado: (2025)
por: Yang, Gao, et al.
Publicado: (2025)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
por: Tong, Zhao, et al.
Publicado: (2026)
por: Tong, Zhao, et al.
Publicado: (2026)
The Curious Case of Curiosity across Human Cultures and LLMs
por: Borah, Angana, et al.
Publicado: (2025)
por: Borah, Angana, et al.
Publicado: (2025)
Do Retrieval Augmented Language Models Know When They Don't Know?
por: Zhou, Youchao, et al.
Publicado: (2025)
por: Zhou, Youchao, et al.
Publicado: (2025)
Deterministic Reversible Data Augmentation for Neural Machine Translation
por: Yao, Jiashu, et al.
Publicado: (2024)
por: Yao, Jiashu, et al.
Publicado: (2024)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
por: Agrawal, Tanmay
Publicado: (2025)
por: Agrawal, Tanmay
Publicado: (2025)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
por: Ma, Zi-Ao, et al.
Publicado: (2025)
por: Ma, Zi-Ao, et al.
Publicado: (2025)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
por: Fu, Yao, et al.
Publicado: (2025)
por: Fu, Yao, et al.
Publicado: (2025)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
por: Ma, Zi-Ao, et al.
Publicado: (2024)
por: Ma, Zi-Ao, et al.
Publicado: (2024)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
por: Sun, Lin, et al.
Publicado: (2025)
por: Sun, Lin, et al.
Publicado: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
por: Zhou, Xiaofeng, et al.
Publicado: (2025)
por: Zhou, Xiaofeng, et al.
Publicado: (2025)
Balancing Stylization and Truth via Disentangled Representation Steering
por: Shen, Chenglei, et al.
Publicado: (2025)
por: Shen, Chenglei, et al.
Publicado: (2025)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
por: Matys, Piotr, et al.
Publicado: (2025)
por: Matys, Piotr, et al.
Publicado: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
por: Kassem, Aly M., et al.
Publicado: (2025)
por: Kassem, Aly M., et al.
Publicado: (2025)
Testing the Limits of Truth Directions in LLMs
por: Poulis, Angelos, et al.
Publicado: (2026)
por: Poulis, Angelos, et al.
Publicado: (2026)
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling
por: Yang, Yizhe, et al.
Publicado: (2022)
por: Yang, Yizhe, et al.
Publicado: (2022)
Word Matters: What Influences Domain Adaptation in Summarization?
por: Li, Yinghao, et al.
Publicado: (2024)
por: Li, Yinghao, et al.
Publicado: (2024)
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
por: Bao, Yuntai, et al.
Publicado: (2025)
por: Bao, Yuntai, et al.
Publicado: (2025)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
por: Adarsh, Shivam, et al.
Publicado: (2026)
por: Adarsh, Shivam, et al.
Publicado: (2026)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
por: Qiu, Tianyi Alex, et al.
Publicado: (2026)
por: Qiu, Tianyi Alex, et al.
Publicado: (2026)
Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
por: Wan, Herun, et al.
Publicado: (2025)
por: Wan, Herun, et al.
Publicado: (2025)
How Good Are LLMs at Out-of-Distribution Detection?
por: Liu, Bo, et al.
Publicado: (2023)
por: Liu, Bo, et al.
Publicado: (2023)
Mix-Initiative Response Generation with Dynamic Prefix Tuning
por: Nie, Yuxiang, et al.
Publicado: (2024)
por: Nie, Yuxiang, et al.
Publicado: (2024)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
por: Reif, Yuval, et al.
Publicado: (2025)
por: Reif, Yuval, et al.
Publicado: (2025)
Training Language Models to Critique With Multi-agent Feedback
por: Lan, Tian, et al.
Publicado: (2024)
por: Lan, Tian, et al.
Publicado: (2024)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
por: Harrasse, Abir, et al.
Publicado: (2025)
por: Harrasse, Abir, et al.
Publicado: (2025)
Ejemplares similares
-
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
por: Liu, Runheng, et al.
Publicado: (2026) -
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
por: Liu, Runheng, et al.
Publicado: (2024) -
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
por: Xiao, Xingchen, et al.
Publicado: (2026) -
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
por: Zhou, Youchao, et al.
Publicado: (2024) -
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
por: Zhang, Guo-Biao, et al.
Publicado: (2026)