The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Christina, Gallagher, Jack, Michala, Jonathan, Fish, Kyle, Lindsey, Jack |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emergent Introspective Awareness in Large Language Models
di: Lindsey, Jack
Pubblicazione: (2026)
di: Lindsey, Jack
Pubblicazione: (2026)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
di: Chen, Runjin, et al.
Pubblicazione: (2025)
di: Chen, Runjin, et al.
Pubblicazione: (2025)
Beyond Static Personas: Situational Personality Steering for Large Language Models
di: Wei, Zesheng, et al.
Pubblicazione: (2026)
di: Wei, Zesheng, et al.
Pubblicazione: (2026)
Improving Agent Interactions in Virtual Environments with Language Models
di: Zhang, Jack
Pubblicazione: (2024)
di: Zhang, Jack
Pubblicazione: (2024)
Slot Machines: How LLMs Keep Track of Multiple Entities
di: Bogdan, Paul C., et al.
Pubblicazione: (2026)
di: Bogdan, Paul C., et al.
Pubblicazione: (2026)
Emotion Concepts and their Function in a Large Language Model
di: Sofroniew, Nicholas, et al.
Pubblicazione: (2026)
di: Sofroniew, Nicholas, et al.
Pubblicazione: (2026)
Aligning VLM Assistants with Personalized Situated Cognition
di: Li, Yongqi, et al.
Pubblicazione: (2025)
di: Li, Yongqi, et al.
Pubblicazione: (2025)
Generics and Default Reasoning in Large Language Models
di: Kirkpatrick, James Ravi, et al.
Pubblicazione: (2025)
di: Kirkpatrick, James Ravi, et al.
Pubblicazione: (2025)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
di: Jordan, Jonathan, et al.
Pubblicazione: (2025)
di: Jordan, Jonathan, et al.
Pubblicazione: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
di: Kim, Geewook, et al.
Pubblicazione: (2024)
di: Kim, Geewook, et al.
Pubblicazione: (2024)
ALMs: Authorial Language Models for Authorship Attribution
di: Huang, Weihang, et al.
Pubblicazione: (2024)
di: Huang, Weihang, et al.
Pubblicazione: (2024)
A Statistical Physics of Language Model Reasoning
di: Carson, Jack David, et al.
Pubblicazione: (2025)
di: Carson, Jack David, et al.
Pubblicazione: (2025)
Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language
di: Alnajjar, Khalid, et al.
Pubblicazione: (2024)
di: Alnajjar, Khalid, et al.
Pubblicazione: (2024)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
di: Young, Jack
Pubblicazione: (2026)
di: Young, Jack
Pubblicazione: (2026)
Persona Jailbreaking in Large Language Models
di: Sandhan, Jivnesh, et al.
Pubblicazione: (2026)
di: Sandhan, Jivnesh, et al.
Pubblicazione: (2026)
Circuit Component Reuse Across Tasks in Transformer Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2023)
di: Merullo, Jack, et al.
Pubblicazione: (2023)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2024)
di: Merullo, Jack, et al.
Pubblicazione: (2024)
Mechanistic Decomposition of Sentence Representations
di: Tehenan, Matthieu, et al.
Pubblicazione: (2025)
di: Tehenan, Matthieu, et al.
Pubblicazione: (2025)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
di: Merullo, Jack, et al.
Pubblicazione: (2023)
di: Merullo, Jack, et al.
Pubblicazione: (2023)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
di: Abdullahi, Tassallah, et al.
Pubblicazione: (2026)
di: Abdullahi, Tassallah, et al.
Pubblicazione: (2026)
Transferring Linear Features Across Language Models With Model Stitching
di: Chen, Alan, et al.
Pubblicazione: (2025)
di: Chen, Alan, et al.
Pubblicazione: (2025)
Situated Natural Language Explanations
di: Zhu, Zining, et al.
Pubblicazione: (2023)
di: Zhu, Zining, et al.
Pubblicazione: (2023)
Large Language Model based Situational Dialogues for Second Language Learning
di: Xu, Shuyao, et al.
Pubblicazione: (2024)
di: Xu, Shuyao, et al.
Pubblicazione: (2024)
Can Large Language Models abstract Medical Coded Language?
di: Lee, Simon A., et al.
Pubblicazione: (2024)
di: Lee, Simon A., et al.
Pubblicazione: (2024)
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
di: Gallagher, Daniel, et al.
Pubblicazione: (2026)
di: Gallagher, Daniel, et al.
Pubblicazione: (2026)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
di: Thomas, Marshall, et al.
Pubblicazione: (2025)
di: Thomas, Marshall, et al.
Pubblicazione: (2025)
EHRmonize: A Framework for Medical Concept Abstraction from Electronic Health Records using Large Language Models
di: Matos, João, et al.
Pubblicazione: (2024)
di: Matos, João, et al.
Pubblicazione: (2024)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
Evaluating Large Language Model Biases in Persona-Steered Generation
di: Liu, Andy, et al.
Pubblicazione: (2024)
di: Liu, Andy, et al.
Pubblicazione: (2024)
Large Language Models as Misleading Assistants in Conversation
di: Hou, Betty Li, et al.
Pubblicazione: (2024)
di: Hou, Betty Li, et al.
Pubblicazione: (2024)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
On Linear Representations and Pretraining Data Frequency in Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2025)
di: Merullo, Jack, et al.
Pubblicazione: (2025)
CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model
di: Lei, Yang, et al.
Pubblicazione: (2023)
di: Lei, Yang, et al.
Pubblicazione: (2023)
Attribute or Abstain: Large Language Models as Long Document Assistants
di: Buchmann, Jan, et al.
Pubblicazione: (2024)
di: Buchmann, Jan, et al.
Pubblicazione: (2024)
A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
di: Decostanzi, Ivan, et al.
Pubblicazione: (2025)
di: Decostanzi, Ivan, et al.
Pubblicazione: (2025)
Probing Persona-Dependent Preferences in Language Models
di: Gilg, Oscar, et al.
Pubblicazione: (2026)
di: Gilg, Oscar, et al.
Pubblicazione: (2026)
Mixture-of-Personas Language Models for Population Simulation
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Emergent Introspective Awareness in Large Language Models
di: Lindsey, Jack
Pubblicazione: (2026) -
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
di: Chen, Runjin, et al.
Pubblicazione: (2025) -
Beyond Static Personas: Situational Personality Steering for Large Language Models
di: Wei, Zesheng, et al.
Pubblicazione: (2026) -
Improving Agent Interactions in Virtual Environments with Language Models
di: Zhang, Jack
Pubblicazione: (2024) -
Slot Machines: How LLMs Keep Track of Multiple Entities
di: Bogdan, Paul C., et al.
Pubblicazione: (2026)