When Truthful Representations Flip Under Deceptive Instructions?
Fuente:
arXiv
Salvato in:
| Autori principali: | Long, Xianxuan, Fu, Yao, Li, Runchao, Sheng, Mu, Yu, Haotian, Han, Xiaotian, Li, Pan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
di: Fu, Yao, et al.
Pubblicazione: (2025)
di: Fu, Yao, et al.
Pubblicazione: (2025)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
di: Fu, Yao, et al.
Pubblicazione: (2025)
di: Fu, Yao, et al.
Pubblicazione: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
di: Wang, Shouren, et al.
Pubblicazione: (2025)
di: Wang, Shouren, et al.
Pubblicazione: (2025)
FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression
di: Li, Runchao, et al.
Pubblicazione: (2025)
di: Li, Runchao, et al.
Pubblicazione: (2025)
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models
di: Fu, Yao, et al.
Pubblicazione: (2024)
di: Fu, Yao, et al.
Pubblicazione: (2024)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
Building Better Deception Probes Using Targeted Instruction Pairs
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
DecepChain: Inducing Deceptive Reasoning in Large Language Models
di: Shen, Wei, et al.
Pubblicazione: (2025)
di: Shen, Wei, et al.
Pubblicazione: (2025)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
di: Wong, Zhen Hao, et al.
Pubblicazione: (2025)
di: Wong, Zhen Hao, et al.
Pubblicazione: (2025)
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
di: Wang, Shouren, et al.
Pubblicazione: (2026)
di: Wang, Shouren, et al.
Pubblicazione: (2026)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
di: Wang, Kai, et al.
Pubblicazione: (2025)
di: Wang, Kai, et al.
Pubblicazione: (2025)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
Spectral Representation-based Reinforcement Learning
di: Gao, Chenxiao, et al.
Pubblicazione: (2025)
di: Gao, Chenxiao, et al.
Pubblicazione: (2025)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips
di: Bacellar, Alan T. L., et al.
Pubblicazione: (2026)
di: Bacellar, Alan T. L., et al.
Pubblicazione: (2026)
Value of Information-based Deceptive Path Planning Under Adversarial Interventions
di: Suttle, Wesley A., et al.
Pubblicazione: (2025)
di: Suttle, Wesley A., et al.
Pubblicazione: (2025)
Relabeling Minimal Training Subset to Flip a Prediction
di: Yang, Jinghan, et al.
Pubblicazione: (2023)
di: Yang, Jinghan, et al.
Pubblicazione: (2023)
FLAIN: Mitigating Backdoor Attacks in Federated Learning via Flipping Weight Updates of Low-Activation Input Neurons
di: Ding, Binbin, et al.
Pubblicazione: (2024)
di: Ding, Binbin, et al.
Pubblicazione: (2024)
HAT-CL: A Hard-Attention-to-the-Task PyTorch Library for Continual Learning
di: Duan, Xiaotian
Pubblicazione: (2023)
di: Duan, Xiaotian
Pubblicazione: (2023)
Deceptive Exploration in Multi-armed Bandits
di: Vurankaya, I. Arda, et al.
Pubblicazione: (2025)
di: Vurankaya, I. Arda, et al.
Pubblicazione: (2025)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
di: Denisov-Blanch, Yegor, et al.
Pubblicazione: (2026)
di: Denisov-Blanch, Yegor, et al.
Pubblicazione: (2026)
Generative Representational Instruction Tuning
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
di: Yang, Wang, et al.
Pubblicazione: (2026)
di: Yang, Wang, et al.
Pubblicazione: (2026)
Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning
di: Cao, Haifang, et al.
Pubblicazione: (2026)
di: Cao, Haifang, et al.
Pubblicazione: (2026)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Information-theoretic Distinctions Between Deception and Confusion
di: Young, Robin
Pubblicazione: (2025)
di: Young, Robin
Pubblicazione: (2025)
MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms
di: Wu, Jinqi, et al.
Pubblicazione: (2026)
di: Wu, Jinqi, et al.
Pubblicazione: (2026)
Representation Invariance and Allocation: When Subgroup Balance Matters
di: Alloula, Anissa, et al.
Pubblicazione: (2025)
di: Alloula, Anissa, et al.
Pubblicazione: (2025)
DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
di: Zhang, Wenxin, et al.
Pubblicazione: (2025)
di: Zhang, Wenxin, et al.
Pubblicazione: (2025)
MergeIT: From Selection to Merging for Efficient Instruction Tuning
di: Cai, Hongyi, et al.
Pubblicazione: (2025)
di: Cai, Hongyi, et al.
Pubblicazione: (2025)
IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
di: Zou, Xiandong, et al.
Pubblicazione: (2025)
di: Zou, Xiandong, et al.
Pubblicazione: (2025)
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
PatternKV: Flattening KV Representation Expands Quantization Headroom
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
Federated Continual Instruction Tuning
di: Guo, Haiyang, et al.
Pubblicazione: (2025)
di: Guo, Haiyang, et al.
Pubblicazione: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
di: Wei, Zhepei, et al.
Pubblicazione: (2025)
di: Wei, Zhepei, et al.
Pubblicazione: (2025)
When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
di: Aghaei, Matin, et al.
Pubblicazione: (2025)
di: Aghaei, Matin, et al.
Pubblicazione: (2025)
Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
di: Kariyappa, Sanjay, et al.
Pubblicazione: (2025)
di: Kariyappa, Sanjay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
di: Fu, Yao, et al.
Pubblicazione: (2025) -
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
di: Fu, Yao, et al.
Pubblicazione: (2025) -
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
di: Wang, Shouren, et al.
Pubblicazione: (2025) -
FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression
di: Li, Runchao, et al.
Pubblicazione: (2025) -
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models
di: Fu, Yao, et al.
Pubblicazione: (2024)