Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination
Fuente:
arXiv
Salvato in:
| Autori principali: | Trivedi, Rakshit, Sharma, Kartik, Parkes, David C |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
di: Sharma, Kartik, et al.
Pubblicazione: (2025)
di: Sharma, Kartik, et al.
Pubblicazione: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
di: Gupta, Pankaj, et al.
Pubblicazione: (2026)
di: Gupta, Pankaj, et al.
Pubblicazione: (2026)
Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
di: Chen, Jiawei, et al.
Pubblicazione: (2026)
di: Chen, Jiawei, et al.
Pubblicazione: (2026)
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
di: Yu, Xingrui, et al.
Pubblicazione: (2024)
di: Yu, Xingrui, et al.
Pubblicazione: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
Can Interpretation Predict Behavior on Unseen Data?
di: Li, Victoria R., et al.
Pubblicazione: (2025)
di: Li, Victoria R., et al.
Pubblicazione: (2025)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
di: Wang, Kaiwen, et al.
Pubblicazione: (2024)
di: Wang, Kaiwen, et al.
Pubblicazione: (2024)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
di: Varshney, Prasoon, et al.
Pubblicazione: (2025)
di: Varshney, Prasoon, et al.
Pubblicazione: (2025)
Are Human Conversations Special? A Large Language Model Perspective
di: Jawale, Toshish, et al.
Pubblicazione: (2024)
di: Jawale, Toshish, et al.
Pubblicazione: (2024)
Learning Policy Representations for Steerable Behavior Synthesis
di: Li, Beiming, et al.
Pubblicazione: (2026)
di: Li, Beiming, et al.
Pubblicazione: (2026)
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
di: Shen, Chengzhi, et al.
Pubblicazione: (2026)
di: Shen, Chengzhi, et al.
Pubblicazione: (2026)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
di: Lam, Michelle S., et al.
Pubblicazione: (2024)
di: Lam, Michelle S., et al.
Pubblicazione: (2024)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
On the Existence and Behavior of Secondary Attention Sinks
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
di: Yang, Zhuonan, et al.
Pubblicazione: (2026)
di: Yang, Zhuonan, et al.
Pubblicazione: (2026)
Brain-Inspired Two-Stage Approach: Enhancing Mathematical Reasoning by Imitating Human Thought Processes
di: Chen, Yezeng, et al.
Pubblicazione: (2024)
di: Chen, Yezeng, et al.
Pubblicazione: (2024)
Eliciting Language Model Behaviors with Investigator Agents
di: Li, Xiang Lisa, et al.
Pubblicazione: (2025)
di: Li, Xiang Lisa, et al.
Pubblicazione: (2025)
Language Models Can Predict Their Own Behavior
di: Ashok, Dhananjay, et al.
Pubblicazione: (2025)
di: Ashok, Dhananjay, et al.
Pubblicazione: (2025)
Large Language Models for Travel Behavior Prediction
di: Mo, Baichuan, et al.
Pubblicazione: (2023)
di: Mo, Baichuan, et al.
Pubblicazione: (2023)
Tool Calling is Linearly Readable and Steerable in Language Models
di: Wu, Zekun, et al.
Pubblicazione: (2026)
di: Wu, Zekun, et al.
Pubblicazione: (2026)
Can Brain Signals Reveal Inner Alignment with Human Languages?
di: Han, William, et al.
Pubblicazione: (2022)
di: Han, William, et al.
Pubblicazione: (2022)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
di: Einwiller, Andreas, et al.
Pubblicazione: (2025)
di: Einwiller, Andreas, et al.
Pubblicazione: (2025)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
di: Bozoukov, Matthew, et al.
Pubblicazione: (2025)
di: Bozoukov, Matthew, et al.
Pubblicazione: (2025)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
di: Yang, Hongming, et al.
Pubblicazione: (2025)
di: Yang, Hongming, et al.
Pubblicazione: (2025)
Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
di: Cao, Jie, et al.
Pubblicazione: (2019)
di: Cao, Jie, et al.
Pubblicazione: (2019)
Diversity Boosts AI-Generated Text Detection
di: Basani, Advik Raj, et al.
Pubblicazione: (2025)
di: Basani, Advik Raj, et al.
Pubblicazione: (2025)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
di: Sastre, Ignacio, et al.
Pubblicazione: (2026)
di: Sastre, Ignacio, et al.
Pubblicazione: (2026)
Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
di: King, Jack, et al.
Pubblicazione: (2026)
di: King, Jack, et al.
Pubblicazione: (2026)
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
di: Mooney, James, et al.
Pubblicazione: (2025)
di: Mooney, James, et al.
Pubblicazione: (2025)
Correcting Large Language Model Behavior via Influence Function
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
di: Yang, Diji, et al.
Pubblicazione: (2024)
di: Yang, Diji, et al.
Pubblicazione: (2024)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
di: Sharma, Yash, et al.
Pubblicazione: (2024)
di: Sharma, Yash, et al.
Pubblicazione: (2024)
Documenti analoghi
-
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
di: Sharma, Kartik, et al.
Pubblicazione: (2026) -
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025) -
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
di: Sharma, Kartik, et al.
Pubblicazione: (2025) -
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
di: Nainani, Jatin, et al.
Pubblicazione: (2024) -
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
di: Gupta, Pankaj, et al.
Pubblicazione: (2026)