Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
Fuente:
arXiv
Saved in:
| Main Authors: | Mujtaba, Dena, Hu, Brian, Hoogs, Anthony, Basharat, Arslan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
by: Adams, Jadie, et al.
Published: (2025)
by: Adams, Jadie, et al.
Published: (2025)
ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
by: Ravichandran, Bharadwaj, et al.
Published: (2025)
by: Ravichandran, Bharadwaj, et al.
Published: (2025)
Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain
by: Hu, Brian, et al.
Published: (2024)
by: Hu, Brian, et al.
Published: (2024)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
by: Zhu, Jian-Qiao, et al.
Published: (2025)
by: Zhu, Jian-Qiao, et al.
Published: (2025)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
by: Lv, Hang, et al.
Published: (2025)
by: Lv, Hang, et al.
Published: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
Test-Time Steering for Lossless Text Compression via Weighted Product of Experts
by: Zhang, Qihang, et al.
Published: (2025)
by: Zhang, Qihang, et al.
Published: (2025)
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
by: Hu, Zhiyuan, et al.
Published: (2026)
by: Hu, Zhiyuan, et al.
Published: (2026)
Adaptive Decoding via Test-Time Policy Learning for Self-Improving Generation
by: Bhardwaj, Asmita, et al.
Published: (2026)
by: Bhardwaj, Asmita, et al.
Published: (2026)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
by: Sun, Moran, et al.
Published: (2026)
by: Sun, Moran, et al.
Published: (2026)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
by: Ye, Wengao, et al.
Published: (2025)
by: Ye, Wengao, et al.
Published: (2025)
Agentic Test-Time Scaling for WebAgents
by: Lee, Nicholas, et al.
Published: (2026)
by: Lee, Nicholas, et al.
Published: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Inference-time Alignment via Sparse Junction Steering
by: Hu, Runyi, et al.
Published: (2026)
by: Hu, Runyi, et al.
Published: (2026)
Context Steering: Controllable Personalization at Inference Time
by: He, Jerry Zhi-Yang, et al.
Published: (2024)
by: He, Jerry Zhi-Yang, et al.
Published: (2024)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
by: Varshney, Prasoon, et al.
Published: (2025)
by: Varshney, Prasoon, et al.
Published: (2025)
Aligning Large Language Model Behavior with Human Citation Preferences
by: Ando, Kenichiro, et al.
Published: (2026)
by: Ando, Kenichiro, et al.
Published: (2026)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
by: Leong, Chak Tou, et al.
Published: (2026)
by: Leong, Chak Tou, et al.
Published: (2026)
TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
by: Chen, Yongchao, et al.
Published: (2025)
by: Chen, Yongchao, et al.
Published: (2025)
BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
by: Ou, Litu, et al.
Published: (2025)
by: Ou, Litu, et al.
Published: (2025)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
Program Synthesis via Test-Time Transduction
by: Lee, Kang-il, et al.
Published: (2025)
by: Lee, Kang-il, et al.
Published: (2025)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
The Real, the Better: Aligning Large Language Models with Online Human Behaviors
by: Jiang, Guanying, et al.
Published: (2024)
by: Jiang, Guanying, et al.
Published: (2024)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization
by: Chen, Yurun, et al.
Published: (2025)
by: Chen, Yurun, et al.
Published: (2025)
Sustainable Digitalization of Business with Multi-Agent RAG and LLM
by: Arslan, Muhammad, et al.
Published: (2025)
by: Arslan, Muhammad, et al.
Published: (2025)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition
by: Li, Yupei, et al.
Published: (2025)
by: Li, Yupei, et al.
Published: (2025)
LLMs Reading the Rhythms of Daily Life: Aligned Understanding for Behavior Prediction and Generation
by: Meng, Fanjin, et al.
Published: (2026)
by: Meng, Fanjin, et al.
Published: (2026)
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
by: Zhang, Wenqi, et al.
Published: (2024)
by: Zhang, Wenqi, et al.
Published: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
Similar Items
-
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
by: Adams, Jadie, et al.
Published: (2025) -
ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
by: Ravichandran, Bharadwaj, et al.
Published: (2025) -
Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain
by: Hu, Brian, et al.
Published: (2024) -
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025) -
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
by: Zhu, Jian-Qiao, et al.
Published: (2025)