Steerability of Instrumental-Convergence Tendencies in LLMs
Fuente:
arXiv
Saved in:
| Main Author: | Hoscilowicz, Jakub |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
by: Hoscilowicz, Jakub, et al.
Published: (2025)
by: Hoscilowicz, Jakub, et al.
Published: (2025)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
Large Language Models for Expansion of Spoken Language Understanding Systems to New Languages
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
Large Language Models as Carriers of Hidden Messages
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
by: Uluoglakci, Cem, et al.
Published: (2024)
by: Uluoglakci, Cem, et al.
Published: (2024)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
by: Chang, Trenton, et al.
Published: (2025)
by: Chang, Trenton, et al.
Published: (2025)
Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes
by: Myers, Quintin, et al.
Published: (2025)
by: Myers, Quintin, et al.
Published: (2025)
CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
by: Nie, Yixin, et al.
Published: (2026)
by: Nie, Yixin, et al.
Published: (2026)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
by: Chen, Zhuomin, et al.
Published: (2025)
by: Chen, Zhuomin, et al.
Published: (2025)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
by: Li, Jing-Jing, et al.
Published: (2024)
by: Li, Jing-Jing, et al.
Published: (2024)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
by: Zhang, Yunfan, et al.
Published: (2025)
by: Zhang, Yunfan, et al.
Published: (2025)
Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs
by: Podolak, Jakub, et al.
Published: (2025)
by: Podolak, Jakub, et al.
Published: (2025)
Bayesian Preference Learning for Test-Time Steerable Reward Models
by: Hong, Jiwoo, et al.
Published: (2026)
by: Hong, Jiwoo, et al.
Published: (2026)
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation
by: Meng, Qingyu, et al.
Published: (2026)
by: Meng, Qingyu, et al.
Published: (2026)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
by: Sorensen, Taylor, et al.
Published: (2025)
by: Sorensen, Taylor, et al.
Published: (2025)
SEAL: Steerable Reasoning Calibration of Large Language Models for Free
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
by: Lee, Jinsook, et al.
Published: (2025)
by: Lee, Jinsook, et al.
Published: (2025)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
by: Huang, Jen-tse, et al.
Published: (2026)
by: Huang, Jen-tse, et al.
Published: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
by: Miehling, Erik, et al.
Published: (2026)
by: Miehling, Erik, et al.
Published: (2026)
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
by: Raju, Prashant C.
Published: (2026)
by: Raju, Prashant C.
Published: (2026)
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
by: Adams, Jadie, et al.
Published: (2025)
by: Adams, Jadie, et al.
Published: (2025)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
by: Wang, Yujun, et al.
Published: (2025)
by: Wang, Yujun, et al.
Published: (2025)
Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
by: Šmíd, Jakub, et al.
Published: (2025)
by: Šmíd, Jakub, et al.
Published: (2025)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
by: Mao, Nathan, et al.
Published: (2026)
by: Mao, Nathan, et al.
Published: (2026)
Tool Calling is Linearly Readable and Steerable in Language Models
by: Wu, Zekun, et al.
Published: (2026)
by: Wu, Zekun, et al.
Published: (2026)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
by: Bigoulaeva, Irina, et al.
Published: (2025)
by: Bigoulaeva, Irina, et al.
Published: (2025)
Sign Language Recognition in the Age of LLMs
by: Javorek, Vaclav, et al.
Published: (2026)
by: Javorek, Vaclav, et al.
Published: (2026)
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
by: Maya, Cristian Espinal
Published: (2026)
by: Maya, Cristian Espinal
Published: (2026)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
by: Macko, Dominik, et al.
Published: (2026)
by: Macko, Dominik, et al.
Published: (2026)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
by: Järviniemi, Olli, et al.
Published: (2024)
by: Järviniemi, Olli, et al.
Published: (2024)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
by: Zhou, Joyce, et al.
Published: (2026)
by: Zhou, Joyce, et al.
Published: (2026)
Similar Items
-
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
by: Hoscilowicz, Jakub, et al.
Published: (2025) -
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023) -
Large Language Models for Expansion of Spoken Language Understanding Systems to New Languages
by: Hoscilowicz, Jakub, et al.
Published: (2024) -
Large Language Models as Carriers of Hidden Messages
by: Hoscilowicz, Jakub, et al.
Published: (2024) -
HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
by: Uluoglakci, Cem, et al.
Published: (2024)