Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Weixu, Yuan, Ye, Han, Changjiang, Tian, Yuxing, Sun, Zipeng, Du, Linfeng, Kang, Jikun, Kang, Hong, Liu, Xue, Wu, Haolun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Message Passing: A Semantic View of Agent Communication Protocols
di: Yuan, Dun, et al.
Pubblicazione: (2026)
di: Yuan, Dun, et al.
Pubblicazione: (2026)
Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization
di: Du, Linfeng, et al.
Pubblicazione: (2026)
di: Du, Linfeng, et al.
Pubblicazione: (2026)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization
di: Yang, Yonghan, et al.
Pubblicazione: (2026)
di: Yang, Yonghan, et al.
Pubblicazione: (2026)
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
di: Zhang, Weixu, et al.
Pubblicazione: (2026)
di: Zhang, Weixu, et al.
Pubblicazione: (2026)
Training Diffusion Language Models for Black-Box Optimization
di: Sun, Zipeng, et al.
Pubblicazione: (2026)
di: Sun, Zipeng, et al.
Pubblicazione: (2026)
QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
di: Yuan, Ye, et al.
Pubblicazione: (2026)
di: Yuan, Ye, et al.
Pubblicazione: (2026)
Diffusion Large Language Models for Black-Box Optimization
di: Yuan, Ye, et al.
Pubblicazione: (2026)
di: Yuan, Ye, et al.
Pubblicazione: (2026)
LLM Safety From Within: Detecting Harmful Content with Internal Representations
di: Jiao, Difan, et al.
Pubblicazione: (2026)
di: Jiao, Difan, et al.
Pubblicazione: (2026)
Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
di: Men, Tianyi, et al.
Pubblicazione: (2024)
di: Men, Tianyi, et al.
Pubblicazione: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
Achieving Tunable Amplified Spontaneous Emission in Rb‐Cs Alloyed Quasi‐2D Perovskites with Low Threshold and Exceptional Spectral Stability
di: Changjiang Shi, et al.
Pubblicazione: (2025)
di: Changjiang Shi, et al.
Pubblicazione: (2025)
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
di: Kim, Jieyong, et al.
Pubblicazione: (2024)
di: Kim, Jieyong, et al.
Pubblicazione: (2024)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
di: Zhang, Qiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Qiyuan, et al.
Pubblicazione: (2025)
ReAttn: Improving Attention-based Re-ranking via Attention Re-weighting
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
A Survey on Data Contamination for Large Language Models
di: Cheng, Yuxing, et al.
Pubblicazione: (2025)
di: Cheng, Yuxing, et al.
Pubblicazione: (2025)
Think Before You Act: Decision Transformers with Working Memory
di: Kang, Jikun, et al.
Pubblicazione: (2023)
di: Kang, Jikun, et al.
Pubblicazione: (2023)
Geospatial Mechanistic Interpretability of Large Language Models
di: De Sabbata, Stef, et al.
Pubblicazione: (2025)
di: De Sabbata, Stef, et al.
Pubblicazione: (2025)
Privis: Towards Content-Aware Secure Volumetric Video Delivery
di: Hu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Hu, Kaiyuan, et al.
Pubblicazione: (2025)
BackSlash: Rate Constrained Optimized Training of Large Language Models
di: Wu, Jun, et al.
Pubblicazione: (2025)
di: Wu, Jun, et al.
Pubblicazione: (2025)
Human Aesthetic Preference-Based Large Text-to-Image Model Personalization: Kandinsky Generation as an Example
di: Zhou, Aven-Le, et al.
Pubblicazione: (2024)
di: Zhou, Aven-Le, et al.
Pubblicazione: (2024)
Learning to Extract Structured Entities Using Language Models
di: Wu, Haolun, et al.
Pubblicazione: (2024)
di: Wu, Haolun, et al.
Pubblicazione: (2024)
Personalized LLM Decoding via Contrasting Personal Preference
di: Bu, Hyungjune, et al.
Pubblicazione: (2025)
di: Bu, Hyungjune, et al.
Pubblicazione: (2025)
SSRLBot: Designing and Developing a Large Language Model-based Agent using Socially Shared Regulated Learning
di: Huang, Xiaoshan, et al.
Pubblicazione: (2025)
di: Huang, Xiaoshan, et al.
Pubblicazione: (2025)
Multivariable $(φ_q,\mathcal{O}_K^{\times})$-modules associated to $p$-adic representations of $\mathrm{Gal}(\overline{K}/K)$
di: Du, Changjiang
Pubblicazione: (2025)
di: Du, Changjiang
Pubblicazione: (2025)
The overconvergence of multivariable $(φ_q,\mathcal{O}_K^{\times})$-modules at the perfectoid level
di: Du, Changjiang
Pubblicazione: (2026)
di: Du, Changjiang
Pubblicazione: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
di: Tak, Ala N., et al.
Pubblicazione: (2025)
di: Tak, Ala N., et al.
Pubblicazione: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
di: Che, Liwei, et al.
Pubblicazione: (2026)
di: Che, Liwei, et al.
Pubblicazione: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
AFSPP: Agent Framework for Shaping Preference and Personality with Large Language Models
di: He, Zihong, et al.
Pubblicazione: (2024)
di: He, Zihong, et al.
Pubblicazione: (2024)
Understanding 6G through Language Models: A Case Study on LLM-aided Structured Entity Extraction in Telecom Domain
di: Yuan, Ye, et al.
Pubblicazione: (2025)
di: Yuan, Ye, et al.
Pubblicazione: (2025)
Learning Discriminative and Generalizable Anomaly Detector for Dynamic Graph with Limited Supervision
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
di: Park, Ji-jun, et al.
Pubblicazione: (2026)
di: Park, Ji-jun, et al.
Pubblicazione: (2026)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
IIMedGPT: Promoting Large Language Model Capabilities of Medical Tasks by Efficient Human Preference Alignment
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
ViRAC: A Vision-Reasoning Agent Head Movement Control Framework in Arbitrary Virtual Environments
di: Hwang, Juyeong, et al.
Pubblicazione: (2025)
di: Hwang, Juyeong, et al.
Pubblicazione: (2025)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
Can Large Language Models Understand Preferences in Personalized Recommendation?
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Message Passing: A Semantic View of Agent Communication Protocols
di: Yuan, Dun, et al.
Pubblicazione: (2026) -
Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization
di: Du, Linfeng, et al.
Pubblicazione: (2026) -
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
di: Tian, Yuxing, et al.
Pubblicazione: (2026) -
Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization
di: Yang, Yonghan, et al.
Pubblicazione: (2026) -
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
di: Zhang, Weixu, et al.
Pubblicazione: (2026)