Steering Language Models Before They Speak: Logit-Level Interventions
Fuente:
arXiv
Guardado en:
| Autores principales: | An, Hyeseon, Park, Shinwoo, Jin, Hyundong, Han, Yo-Sub |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
Linguistics-Aware Non-Distortionary LLM Watermarking
por: Park, Shinwoo, et al.
Publicado: (2026)
por: Park, Shinwoo, et al.
Publicado: (2026)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
por: Jin, Hyundong, et al.
Publicado: (2026)
por: Jin, Hyundong, et al.
Publicado: (2026)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
por: Park, Shinwoo, et al.
Publicado: (2026)
por: Park, Shinwoo, et al.
Publicado: (2026)
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
por: Jin, Hyundong, et al.
Publicado: (2026)
por: Jin, Hyundong, et al.
Publicado: (2026)
Sequential Behavioral Watermarking for LLM Agents
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
por: An, Hyeseon, et al.
Publicado: (2025)
por: An, Hyeseon, et al.
Publicado: (2025)
WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents
por: Jin, Hyundong, et al.
Publicado: (2025)
por: Jin, Hyundong, et al.
Publicado: (2025)
Detection of LLM-Paraphrased Code and Identification of the Responsible LLM Using Coding Style Features
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
por: Jin, Hyundong, et al.
Publicado: (2025)
por: Jin, Hyundong, et al.
Publicado: (2025)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
por: Kim, Jungin, et al.
Publicado: (2025)
por: Kim, Jungin, et al.
Publicado: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
por: Lee, Yejin, et al.
Publicado: (2026)
por: Lee, Yejin, et al.
Publicado: (2026)
Thinking Before Speaking: A Role-playing Model with Mindset
por: Zhang, Baohua, et al.
Publicado: (2024)
por: Zhang, Baohua, et al.
Publicado: (2024)
How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
The Information Geometry of Softmax: Probing and Steering
por: Park, Kiho, et al.
Publicado: (2026)
por: Park, Kiho, et al.
Publicado: (2026)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
por: Wang, Xintong, et al.
Publicado: (2024)
por: Wang, Xintong, et al.
Publicado: (2024)
Think Before You Speak: Cultivating Communication Skills of Large Language Models via Inner Monologue
por: Zhou, Junkai, et al.
Publicado: (2023)
por: Zhou, Junkai, et al.
Publicado: (2023)
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
por: Zelikman, Eric, et al.
Publicado: (2024)
por: Zelikman, Eric, et al.
Publicado: (2024)
STAB: Specification-driven Testing for Algorithmic Bottlenecks
por: Lim, Soohan, et al.
Publicado: (2026)
por: Lim, Soohan, et al.
Publicado: (2026)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
por: Kang, Diancheng, et al.
Publicado: (2026)
por: Kang, Diancheng, et al.
Publicado: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
por: Park, Kiho, et al.
Publicado: (2023)
por: Park, Kiho, et al.
Publicado: (2023)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
por: Li, Jin, et al.
Publicado: (2025)
por: Li, Jin, et al.
Publicado: (2025)
SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
por: Acikgoz, Emre Can, et al.
Publicado: (2025)
por: Acikgoz, Emre Can, et al.
Publicado: (2025)
Self-Steering Language Models
por: Grand, Gabriel, et al.
Publicado: (2025)
por: Grand, Gabriel, et al.
Publicado: (2025)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
por: Zhou, Hanhan, et al.
Publicado: (2026)
por: Zhou, Hanhan, et al.
Publicado: (2026)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
por: Cheng, Zifeng, et al.
Publicado: (2025)
por: Cheng, Zifeng, et al.
Publicado: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
por: Park, Kiho, et al.
Publicado: (2024)
por: Park, Kiho, et al.
Publicado: (2024)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
por: Zeng, Jingying, et al.
Publicado: (2025)
por: Zeng, Jingying, et al.
Publicado: (2025)
UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models
por: Dong, Yijiang River, et al.
Publicado: (2024)
por: Dong, Yijiang River, et al.
Publicado: (2024)
On the Limitations of Steering in Language Model Alignment
por: Niranjan, Chebrolu, et al.
Publicado: (2025)
por: Niranjan, Chebrolu, et al.
Publicado: (2025)
Word Embeddings Are Steers for Language Models
por: Han, Chi, et al.
Publicado: (2023)
por: Han, Chi, et al.
Publicado: (2023)
Ejemplares similares
-
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
por: An, Hyeseon, et al.
Publicado: (2026) -
Linguistics-Aware Non-Distortionary LLM Watermarking
por: Park, Shinwoo, et al.
Publicado: (2026) -
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
por: Park, Shinwoo, et al.
Publicado: (2025) -
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
por: Jin, Hyundong, et al.
Publicado: (2026) -
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
por: Park, Shinwoo, et al.
Publicado: (2026)