Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Tianlong, Jiao, Xianfeng, Zhu, Yinghao, Chen, Zhongzhi, He, Yifan, Chu, Xu, Gao, Junyi, Wang, Yasha, Ma, Liantao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
di: Liao, Weibin, et al.
Pubblicazione: (2025)
di: Liao, Weibin, et al.
Pubblicazione: (2025)
ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis
di: Zhao, Huiya, et al.
Pubblicazione: (2025)
di: Zhao, Huiya, et al.
Pubblicazione: (2025)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
di: Chen, Zhongzhi, et al.
Pubblicazione: (2023)
di: Chen, Zhongzhi, et al.
Pubblicazione: (2023)
A Comprehensive Benchmark for COVID-19 Predictive Modeling Using Electronic Health Records in Intensive Care
di: Gao, Junyi, et al.
Pubblicazione: (2022)
di: Gao, Junyi, et al.
Pubblicazione: (2022)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent Collaboration
di: Wang, Zixiang, et al.
Pubblicazione: (2024)
di: Wang, Zixiang, et al.
Pubblicazione: (2024)
LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation
di: Liao, Weibin, et al.
Pubblicazione: (2024)
di: Liao, Weibin, et al.
Pubblicazione: (2024)
Learnable Prompt as Pseudo-Imputation: Rethinking the Necessity of Traditional EHR Data Imputation in Downstream Clinical Prediction
di: Liao, Weibin, et al.
Pubblicazione: (2024)
di: Liao, Weibin, et al.
Pubblicazione: (2024)
Domain-invariant Clinical Representation Learning by Bridging Data Distribution Shift across EMR Datasets
di: Zhang, Zhongji, et al.
Pubblicazione: (2023)
di: Zhang, Zhongji, et al.
Pubblicazione: (2023)
Imputation with Inter-Series Information from Prototypes for Irregular Sampled Time Series
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
di: Luo, Wen, et al.
Pubblicazione: (2026)
di: Luo, Wen, et al.
Pubblicazione: (2026)
RankSteer: Activation Steering for Pointwise LLM Ranking
di: Wang, Yumeng, et al.
Pubblicazione: (2026)
di: Wang, Yumeng, et al.
Pubblicazione: (2026)
Auditing medical multi-agent AI reveals risks of false consensus
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and Obstetrics
di: Zhu, Yinghao, et al.
Pubblicazione: (2026)
di: Zhu, Yinghao, et al.
Pubblicazione: (2026)
Prompting Large Language Models for Zero-Shot Clinical Prediction with Structured Longitudinal Electronic Health Record Data
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning
di: Zhang, Jinyang, et al.
Pubblicazione: (2025)
di: Zhang, Jinyang, et al.
Pubblicazione: (2025)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
di: Tan, Zhen, et al.
Pubblicazione: (2024)
di: Tan, Zhen, et al.
Pubblicazione: (2024)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
di: Wannan, et al.
Pubblicazione: (2025)
di: Wannan, et al.
Pubblicazione: (2025)
Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
di: Fang, Yue, et al.
Pubblicazione: (2025)
di: Fang, Yue, et al.
Pubblicazione: (2025)
The intellectual base and research fronts of IL‐18: A bibliometric review of the literature from WoSCC (2012–2022)
di: Zhongzhi Wang
Pubblicazione: (2024)
di: Zhongzhi Wang
Pubblicazione: (2024)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
PRISM: Mitigating EHR Data Sparsity via Learning from Missing Feature Calibrated Prototype Patient Representations
di: Zhu, Yinghao, et al.
Pubblicazione: (2023)
di: Zhu, Yinghao, et al.
Pubblicazione: (2023)
TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference Trees
di: Liao, Weibin, et al.
Pubblicazione: (2024)
di: Liao, Weibin, et al.
Pubblicazione: (2024)
Dual-Axis Beam-Steering OPA with purely Passive Phase Shifters
di: Kakdarvishi, Venus, et al.
Pubblicazione: (2024)
di: Kakdarvishi, Venus, et al.
Pubblicazione: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
di: O'Neill, Charles, et al.
Pubblicazione: (2025)
di: O'Neill, Charles, et al.
Pubblicazione: (2025)
Steered LLM Activations are Non-Surjective
di: Mishra, Aayush, et al.
Pubblicazione: (2026)
di: Mishra, Aayush, et al.
Pubblicazione: (2026)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
di: Wang, Weixuan, et al.
Pubblicazione: (2024)
di: Wang, Weixuan, et al.
Pubblicazione: (2024)
Test-time Diverse Reasoning by Riemannian Activation Steering
di: Khanh, Ly Tran Ho, et al.
Pubblicazione: (2025)
di: Khanh, Ly Tran Ho, et al.
Pubblicazione: (2025)
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
di: Zou, Zhengtao, et al.
Pubblicazione: (2025)
di: Zou, Zhengtao, et al.
Pubblicazione: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
di: Wang, Guangzhi, et al.
Pubblicazione: (2025)
di: Wang, Guangzhi, et al.
Pubblicazione: (2025)
GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
DecomPose: Disentangling Cross-Category Optimization Contention for Category-Level 6D Object Pose Estimation
di: Gao, Yifan, et al.
Pubblicazione: (2026)
di: Gao, Yifan, et al.
Pubblicazione: (2026)
HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
di: Liao, Weibin, et al.
Pubblicazione: (2025) -
ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis
di: Zhao, Huiya, et al.
Pubblicazione: (2025) -
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
di: Chen, Zhongzhi, et al.
Pubblicazione: (2023) -
A Comprehensive Benchmark for COVID-19 Predictive Modeling Using Electronic Health Records in Intensive Care
di: Gao, Junyi, et al.
Pubblicazione: (2022) -
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
di: Zhu, Yinghao, et al.
Pubblicazione: (2025)