Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
Fuente:
arXiv
Saved in:
| Main Authors: | Banayeeanzade, Amin, Tak, Ala N., Bahrani, Fatemeh, Bolourani, Anahita, Blas, Leonardo, Ferrara, Emilio, Gratch, Jonathan, Karimireddy, Sai Praneeth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026)
by: Tak, Ala N., et al.
Published: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
Psychological Steering of Large Language Models
by: Blas, Leonardo, et al.
Published: (2026)
by: Blas, Leonardo, et al.
Published: (2026)
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
by: Tak, Ala N., et al.
Published: (2024)
by: Tak, Ala N., et al.
Published: (2024)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
by: Choi, Jacob, et al.
Published: (2025)
by: Choi, Jacob, et al.
Published: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
by: Kwon, Deuksin, et al.
Published: (2024)
by: Kwon, Deuksin, et al.
Published: (2024)
A Systematic Analysis of Base Model Choice for Reward Modeling
by: Ahrabian, Kian, et al.
Published: (2025)
by: Ahrabian, Kian, et al.
Published: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
by: Hallinan, Skyler, et al.
Published: (2025)
by: Hallinan, Skyler, et al.
Published: (2025)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
by: Gong, Litian, et al.
Published: (2025)
by: Gong, Litian, et al.
Published: (2025)
Optimization with Access to Auxiliary Information
by: Chayti, El Mahdi, et al.
Published: (2022)
by: Chayti, El Mahdi, et al.
Published: (2022)
OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
by: Hallinan, Skyler, et al.
Published: (2026)
by: Hallinan, Skyler, et al.
Published: (2026)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
by: Ziashahabi, Amir, et al.
Published: (2025)
by: Ziashahabi, Amir, et al.
Published: (2025)
On the Limits of Momentum in Decentralized and Federated Optimization
by: Zaccone, Riccardo, et al.
Published: (2025)
by: Zaccone, Riccardo, et al.
Published: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
LIA: Privacy-Preserving Data Quality Evaluation in Federated Learning Using a Lazy Influence Approximation
by: Rokvic, Ljubomir, et al.
Published: (2022)
by: Rokvic, Ljubomir, et al.
Published: (2022)
Entropy-driven Fair and Effective Federated Learning
by: Wang, Lin, et al.
Published: (2023)
by: Wang, Lin, et al.
Published: (2023)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
by: Bakman, Yavuz, et al.
Published: (2025)
by: Bakman, Yavuz, et al.
Published: (2025)
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
by: Dey, Priyanka, et al.
Published: (2025)
by: Dey, Priyanka, et al.
Published: (2025)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
by: Hasani, Hosein, et al.
Published: (2026)
by: Hasani, Hosein, et al.
Published: (2026)
Personality Expression Across Contexts: Linguistic and Behavioral Variation in LLM Agents
by: Han, Bin, et al.
Published: (2026)
by: Han, Bin, et al.
Published: (2026)
Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
by: Bakman, Yavuz, et al.
Published: (2026)
by: Bakman, Yavuz, et al.
Published: (2026)
Collaborative Heterogeneous Causal Inference Beyond Meta-analysis
by: Guo, Tianyu, et al.
Published: (2024)
by: Guo, Tianyu, et al.
Published: (2024)
Defection-Free Collaboration between Competitors in a Learning System
by: Werner, Mariel, et al.
Published: (2024)
by: Werner, Mariel, et al.
Published: (2024)
Do Data Valuations Make Good Data Prices?
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
Can Language Model Moderators Improve the Health of Online Discourse?
by: Cho, Hyundong, et al.
Published: (2023)
by: Cho, Hyundong, et al.
Published: (2023)
Robust Multi-Agent LLMs under Byzantine Faults
by: Lee, Haejoon, et al.
Published: (2026)
by: Lee, Haejoon, et al.
Published: (2026)
FACT-GPT: Fact-Checking Augmentation via Claim Matching with LLMs
by: Choi, Eun Cheol, et al.
Published: (2024)
by: Choi, Eun Cheol, et al.
Published: (2024)
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
Unearthing a Billion Telegram Posts about the 2024 U.S. Presidential Election: Development of a Public Dataset
by: Blas, Leonardo, et al.
Published: (2024)
by: Blas, Leonardo, et al.
Published: (2024)
GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
by: Ferrara, Emilio
Published: (2023)
by: Ferrara, Emilio
Published: (2023)
Secret Keepers: The Impact of LLMs on Linguistic Markers of Personal Traits
by: Sourati, Zhivar, et al.
Published: (2024)
by: Sourati, Zhivar, et al.
Published: (2024)
Communication-Efficient Heterogeneous Federated Learning with Generalized Heavy-Ball Momentum
by: Zaccone, Riccardo, et al.
Published: (2023)
by: Zaccone, Riccardo, et al.
Published: (2023)
A Differentially Private Kaplan-Meier Estimator for Privacy-Preserving Survival Analysis
by: Veeraragavan, Narasimha Raghavan, et al.
Published: (2024)
by: Veeraragavan, Narasimha Raghavan, et al.
Published: (2024)
DAVED: Data Acquisition via Experimental Design for Data Markets
by: Lu, Charles, et al.
Published: (2024)
by: Lu, Charles, et al.
Published: (2024)
Similar Items
-
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026) -
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025) -
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
by: Banayeeanzade, Amin, et al.
Published: (2026) -
Psychological Steering of Large Language Models
by: Blas, Leonardo, et al.
Published: (2026) -
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
by: Tak, Ala N., et al.
Published: (2024)