Controllable and explainable personality sliders for LLMs at inference time
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hoppe, Florian, Khachaturov, David, Mullins, Robert, Meng, Mark Huasong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Complexity Matters: Effective Dimensionality as a Measure for Adversarial Robustness
par: Khachaturov, David, et autres
Publié: (2024)
par: Khachaturov, David, et autres
Publié: (2024)
Verbalizing LLMs' assumptions to explain and control sycophancy
par: Cheng, Myra, et autres
Publié: (2026)
par: Cheng, Myra, et autres
Publié: (2026)
Feedback Forensics: A Toolkit to Measure AI Personality
par: Findeis, Arduin, et autres
Publié: (2025)
par: Findeis, Arduin, et autres
Publié: (2025)
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
par: Zhang, Zhibo, et autres
Publié: (2024)
par: Zhang, Zhibo, et autres
Publié: (2024)
Inverse Constitutional AI: Compressing Preferences into Principles
par: Findeis, Arduin, et autres
Publié: (2024)
par: Findeis, Arduin, et autres
Publié: (2024)
Watermarking Needs Input Repetition Masking
par: Khachaturov, David, et autres
Publié: (2025)
par: Khachaturov, David, et autres
Publié: (2025)
Large language models struggle with ethnographic text annotation
par: Goodall, Leonardo S., et autres
Publié: (2026)
par: Goodall, Leonardo S., et autres
Publié: (2026)
Enhancing reasoning accuracy in large language models during inference time
par: Sharma, Vinay, et autres
Publié: (2026)
par: Sharma, Vinay, et autres
Publié: (2026)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
par: Yang, Chang, et autres
Publié: (2025)
par: Yang, Chang, et autres
Publié: (2025)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
par: Li, Wenjie, et autres
Publié: (2026)
par: Li, Wenjie, et autres
Publié: (2026)
Learning and Enforcing Context-Sensitive Control for LLMs
par: Albinhassan, Mohammad, et autres
Publié: (2026)
par: Albinhassan, Mohammad, et autres
Publié: (2026)
Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager
par: Galland, Lucie, et autres
Publié: (2025)
par: Galland, Lucie, et autres
Publié: (2025)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
par: Khachaturov, David, et autres
Publié: (2025)
par: Khachaturov, David, et autres
Publié: (2025)
Detecting and explaining postpartum depression in real-time with generative artificial intelligence
par: García-Méndez, Silvia, et autres
Publié: (2025)
par: García-Méndez, Silvia, et autres
Publié: (2025)
Investigating the Robustness of Deductive Reasoning with Large Language Models
par: Hoppe, Fabian, et autres
Publié: (2025)
par: Hoppe, Fabian, et autres
Publié: (2025)
Two-dimensional early exit optimisation of LLM inference
par: Hůla, Jan, et autres
Publié: (2026)
par: Hůla, Jan, et autres
Publié: (2026)
Testing the Limits of Truth Directions in LLMs
par: Poulis, Angelos, et autres
Publié: (2026)
par: Poulis, Angelos, et autres
Publié: (2026)
Fairness Evaluation and Inference Level Mitigation in LLMs
par: Nadeem, Afrozah, et autres
Publié: (2025)
par: Nadeem, Afrozah, et autres
Publié: (2025)
An overview of model uncertainty and variability in LLM-based sentiment analysis. Challenges, mitigation strategies and the role of explainability
par: Herrera-Poyatos, David, et autres
Publié: (2025)
par: Herrera-Poyatos, David, et autres
Publié: (2025)
Multi-Personality Generation of LLMs at Decoding-time
par: Chen, Rongxin, et autres
Publié: (2025)
par: Chen, Rongxin, et autres
Publié: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
par: Nadeem, Afrozah, et autres
Publié: (2025)
par: Nadeem, Afrozah, et autres
Publié: (2025)
KV Cache Steering for Controlling Frozen LLMs
par: Belitsky, Max, et autres
Publié: (2025)
par: Belitsky, Max, et autres
Publié: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
par: Afzal, Anum, et autres
Publié: (2024)
par: Afzal, Anum, et autres
Publié: (2024)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
par: Nadeem, Afrozah, et autres
Publié: (2025)
par: Nadeem, Afrozah, et autres
Publié: (2025)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
par: Wang, Chenxi, et autres
Publié: (2025)
par: Wang, Chenxi, et autres
Publié: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
par: Xia, Heming, et autres
Publié: (2025)
par: Xia, Heming, et autres
Publié: (2025)
CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
par: Xue, Xiaona, et autres
Publié: (2026)
par: Xue, Xiaona, et autres
Publié: (2026)
Using LLMs to identify features of personal and professional skills in an open-response situational judgment test
par: Walsh, Cole, et autres
Publié: (2025)
par: Walsh, Cole, et autres
Publié: (2025)
PAD: Personalized Alignment of LLMs at Decoding-Time
par: Chen, Ruizhe, et autres
Publié: (2024)
par: Chen, Ruizhe, et autres
Publié: (2024)
Can We Edit LLMs for Long-Tail Biomedical Knowledge?
par: Yi, Xinhao, et autres
Publié: (2025)
par: Yi, Xinhao, et autres
Publié: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
par: Arbabi, Alireza, et autres
Publié: (2025)
par: Arbabi, Alireza, et autres
Publié: (2025)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
par: An, Bang, et autres
Publié: (2025)
par: An, Bang, et autres
Publié: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
par: Tan, Daniel, et autres
Publié: (2025)
par: Tan, Daniel, et autres
Publié: (2025)
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
par: Rawat, Shivam, et autres
Publié: (2026)
par: Rawat, Shivam, et autres
Publié: (2026)
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
par: Bannò, Stefano, et autres
Publié: (2025)
par: Bannò, Stefano, et autres
Publié: (2025)
Attribute-Aware Controlled Product Generation with LLMs for E-commerce
par: Negri, Virginia, et autres
Publié: (2025)
par: Negri, Virginia, et autres
Publié: (2025)
Evaluating and explaining training strategies for zero-shot cross-lingual news sentiment analysis
par: Andrenšek, Luka, et autres
Publié: (2024)
par: Andrenšek, Luka, et autres
Publié: (2024)
An explainable transformer circuit for compositional generalization
par: Tang, Cheng, et autres
Publié: (2025)
par: Tang, Cheng, et autres
Publié: (2025)
The Political Preferences of LLMs
par: Rozado, David
Publié: (2024)
par: Rozado, David
Publié: (2024)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
par: Schlangen, David
Publié: (2024)
par: Schlangen, David
Publié: (2024)
Documents similaires
-
Complexity Matters: Effective Dimensionality as a Measure for Adversarial Robustness
par: Khachaturov, David, et autres
Publié: (2024) -
Verbalizing LLMs' assumptions to explain and control sycophancy
par: Cheng, Myra, et autres
Publié: (2026) -
Feedback Forensics: A Toolkit to Measure AI Personality
par: Findeis, Arduin, et autres
Publié: (2025) -
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
par: Zhang, Zhibo, et autres
Publié: (2024) -
Inverse Constitutional AI: Compressing Preferences into Principles
par: Findeis, Arduin, et autres
Publié: (2024)