Effectively Steer LLM To Follow Preference via Building Confident Directions
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Bingqing, Han, Boran, Zhang, Shuai, Wang, Hao, Fang, Haoyang, Min, Bonan, Wang, Yuyang, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025)
by: Adila, Dyah, et al.
Published: (2025)
Efficient Table Retrieval and Understanding with Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2026)
by: Xu, Zhuoyan, et al.
Published: (2026)
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
by: Song, Feifan, et al.
Published: (2024)
by: Song, Feifan, et al.
Published: (2024)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025)
by: Tsaknakis, Ioannis, et al.
Published: (2025)
Unraveling the Gradient Descent Dynamics of Transformers
by: Song, Bingqing, et al.
Published: (2024)
by: Song, Bingqing, et al.
Published: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
HelpSteer2-Preference: Complementing Ratings with Preferences
by: Wang, Zhilin, et al.
Published: (2024)
by: Wang, Zhilin, et al.
Published: (2024)
DROJ: A Prompt-Driven Attack against Large Language Models
by: Hu, Leyang, et al.
Published: (2024)
by: Hu, Leyang, et al.
Published: (2024)
CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaboration
by: Jia, Zhihao, et al.
Published: (2025)
by: Jia, Zhihao, et al.
Published: (2025)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
by: Lv, Hang, et al.
Published: (2025)
by: Lv, Hang, et al.
Published: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
by: Chen, Jiaqi, et al.
Published: (2026)
by: Chen, Jiaqi, et al.
Published: (2026)
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
by: Bao, Qiming, et al.
Published: (2026)
by: Bao, Qiming, et al.
Published: (2026)
Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
by: Heyman, Geert, et al.
Published: (2026)
by: Heyman, Geert, et al.
Published: (2026)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
by: Lu, Junjie, et al.
Published: (2025)
by: Lu, Junjie, et al.
Published: (2025)
Steering LLM Thinking with Budget Guidance
by: Li, Junyan, et al.
Published: (2025)
by: Li, Junyan, et al.
Published: (2025)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
Teaching LLMs to Abstain via Fine-Grained Semantic Confidence Reward
by: An, Hao, et al.
Published: (2025)
by: An, Hao, et al.
Published: (2025)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
by: Cui, Guofeng, et al.
Published: (2025)
by: Cui, Guofeng, et al.
Published: (2025)
SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning
by: He, Zelin, et al.
Published: (2026)
by: He, Zelin, et al.
Published: (2026)
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
by: Afzali, Amirabbas, et al.
Published: (2026)
by: Afzali, Amirabbas, et al.
Published: (2026)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
by: Song, Yiliang, et al.
Published: (2026)
by: Song, Yiliang, et al.
Published: (2026)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Similar Items
-
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025) -
Efficient Table Retrieval and Understanding with Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2026) -
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
by: Song, Feifan, et al.
Published: (2024) -
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025) -
Unraveling the Gradient Descent Dynamics of Transformers
by: Song, Bingqing, et al.
Published: (2024)