Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Kia-Jüng, Meier, Dominik, Zhao, Jiachen, Ruas, Terry, Gipp, Bela |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
What's Wrong? Refining Meeting Summaries with LLM Feedback
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
Tell me what I need to know: Exploring LLM-based (Personalized) Abstractive Multi-Source Meeting Summarization
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
Re-FRAME the Meeting Summarization SCOPE: Fact-Based Summarization and Personalization via Questions
por: Kirstein, Frederic, et al.
Publicado: (2025)
por: Kirstein, Frederic, et al.
Publicado: (2025)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
CADS: A Systematic Literature Review on the Challenges of Abstractive Dialogue Summarization
por: Kirstein, Frederic, et al.
Publicado: (2024)
por: Kirstein, Frederic, et al.
Publicado: (2024)
How Large Language Models are Transforming Machine-Paraphrased Plagiarism
por: Wahle, Jan Philip, et al.
Publicado: (2022)
por: Wahle, Jan Philip, et al.
Publicado: (2022)
SPaRC: A Spatial Pathfinding Reasoning Challenge
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2025)
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2025)
You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with a Multi-Agent Conversations
por: Kirstein, Frederic, et al.
Publicado: (2025)
por: Kirstein, Frederic, et al.
Publicado: (2025)
Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
por: Gnewuch, Max Martin, et al.
Publicado: (2025)
por: Gnewuch, Max Martin, et al.
Publicado: (2025)
ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
por: Yang, Tianyu, et al.
Publicado: (2025)
por: Yang, Tianyu, et al.
Publicado: (2025)
Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling
por: Wunderlich, Florian Valentin, et al.
Publicado: (2026)
por: Wunderlich, Florian Valentin, et al.
Publicado: (2026)
Voting or Consensus? Decision-Making in Multi-Agent Debate
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2025)
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2025)
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2026)
por: Kaesberg, Lars Benedikt, et al.
Publicado: (2026)
Towards Human Understanding of Paraphrase Types in Large Language Models
por: Meier, Dominik, et al.
Publicado: (2024)
por: Meier, Dominik, et al.
Publicado: (2024)
Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
por: Wahle, Jan Philip, et al.
Publicado: (2021)
por: Wahle, Jan Philip, et al.
Publicado: (2021)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
por: Alagharu, Rishab, et al.
Publicado: (2026)
por: Alagharu, Rishab, et al.
Publicado: (2026)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
por: Meier, Dominik, et al.
Publicado: (2025)
por: Meier, Dominik, et al.
Publicado: (2025)
MALLM: Multi-Agent Large Language Models Framework
por: Becker, Jonas, et al.
Publicado: (2025)
por: Becker, Jonas, et al.
Publicado: (2025)
LogSigma at SemEval-2026 Task 3: Uncertainty-Weighted Multitask Learning for Dimensional Aspect-Based Sentiment Analysis
por: Hikal, Baraa, et al.
Publicado: (2026)
por: Hikal, Baraa, et al.
Publicado: (2026)
Activation Steering for Chain-of-Thought Compression
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
Paraphrase Types for Generation and Detection
por: Wahle, Jan Philip, et al.
Publicado: (2023)
por: Wahle, Jan Philip, et al.
Publicado: (2023)
Beyond Single-Granularity Prompts: A Multi-Scale Chain-of-Thought Prompt Learning for Graph
por: Zheng, Ziyu, et al.
Publicado: (2025)
por: Zheng, Ziyu, et al.
Publicado: (2025)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
por: Sheng, Leheng, et al.
Publicado: (2025)
por: Sheng, Leheng, et al.
Publicado: (2025)
Refusal in Language Models Is Mediated by a Single Direction
por: Arditi, Andy, et al.
Publicado: (2024)
por: Arditi, Andy, et al.
Publicado: (2024)
Paraphrase Types Elicit Prompt Engineering Capabilities
por: Wahle, Jan Philip, et al.
Publicado: (2024)
por: Wahle, Jan Philip, et al.
Publicado: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
por: García-Ferrero, Iker, et al.
Publicado: (2025)
por: García-Ferrero, Iker, et al.
Publicado: (2025)
Programming Refusal with Conditional Activation Steering
por: Lee, Bruce W., et al.
Publicado: (2024)
por: Lee, Bruce W., et al.
Publicado: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
por: Yang, Jiaxi, et al.
Publicado: (2026)
por: Yang, Jiaxi, et al.
Publicado: (2026)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
por: Schmidt, Finn, et al.
Publicado: (2026)
por: Schmidt, Finn, et al.
Publicado: (2026)
Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification
por: Zhukova, Anastasia, et al.
Publicado: (2026)
por: Zhukova, Anastasia, et al.
Publicado: (2026)
Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges
por: Becker, Jonas, et al.
Publicado: (2024)
por: Becker, Jonas, et al.
Publicado: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
por: Cheng, Stephen, et al.
Publicado: (2026)
por: Cheng, Stephen, et al.
Publicado: (2026)
Chain-of-Thought Hijacking
por: Zhao, Jianli, et al.
Publicado: (2025)
por: Zhao, Jianli, et al.
Publicado: (2025)
Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
por: Cox, Kyle, et al.
Publicado: (2026)
por: Cox, Kyle, et al.
Publicado: (2026)
LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
por: Shu, Huizhen, et al.
Publicado: (2025)
por: Shu, Huizhen, et al.
Publicado: (2025)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
por: Xia, Yu, et al.
Publicado: (2024)
por: Xia, Yu, et al.
Publicado: (2024)
Chain-of-Thought in Large Language Models: Decoding, Projection, and Activation
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
por: Beel, Joeran, et al.
Publicado: (2025)
por: Beel, Joeran, et al.
Publicado: (2025)
Ejemplares similares
-
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
por: Kirstein, Frederic, et al.
Publicado: (2024) -
What's Wrong? Refining Meeting Summaries with LLM Feedback
por: Kirstein, Frederic, et al.
Publicado: (2024) -
Tell me what I need to know: Exploring LLM-based (Personalized) Abstractive Multi-Source Meeting Summarization
por: Kirstein, Frederic, et al.
Publicado: (2024) -
Re-FRAME the Meeting Summarization SCOPE: Fact-Based Summarization and Personalization via Questions
por: Kirstein, Frederic, et al.
Publicado: (2025) -
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
por: Kirstein, Frederic, et al.
Publicado: (2024)