General and Efficient Steering of Unconditional Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qingsong, Belkin, Mikhail, Wang, Yusu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Out-of-Distribution Detection with a Single Unconditional Diffusion Model
von: Heng, Alvin, et al.
Veröffentlicht: (2024)
von: Heng, Alvin, et al.
Veröffentlicht: (2024)
DiffLoad: Uncertainty Quantification in Electrical Load Forecasting with the Diffusion Model
von: Wang, Zhixian, et al.
Veröffentlicht: (2023)
von: Wang, Zhixian, et al.
Veröffentlicht: (2023)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
MidSteer: Optimal Affine Framework for Steering Generative Models
von: Gaintseva, Tatiana, et al.
Veröffentlicht: (2026)
von: Gaintseva, Tatiana, et al.
Veröffentlicht: (2026)
Task-oriented Time Series Imputation Evaluation via Generalized Representers
von: Wang, Zhixian, et al.
Veröffentlicht: (2024)
von: Wang, Zhixian, et al.
Veröffentlicht: (2024)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
von: He, Jesse, et al.
Veröffentlicht: (2026)
von: He, Jesse, et al.
Veröffentlicht: (2026)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
DGPO: RL-Steered Graph Diffusion for Neural Architecture Generation
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
von: Rector-Brooks, Jarrid, et al.
Veröffentlicht: (2024)
von: Rector-Brooks, Jarrid, et al.
Veröffentlicht: (2024)
Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Prediction
von: Tang, Zhihao, et al.
Veröffentlicht: (2025)
von: Tang, Zhihao, et al.
Veröffentlicht: (2025)
BarrierSteer: LLM Safety via Learning Barrier Steering
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
Generalized Discrete Diffusion with Self-Correction
von: Wang, Linxuan, et al.
Veröffentlicht: (2026)
von: Wang, Linxuan, et al.
Veröffentlicht: (2026)
A Triple-Inertial Accelerated Alternating Optimization Method for Deep Learning Training
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search
von: Robertson, John T., et al.
Veröffentlicht: (2026)
von: Robertson, John T., et al.
Veröffentlicht: (2026)
Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
von: Wang, Qiongyan, et al.
Veröffentlicht: (2025)
von: Wang, Qiongyan, et al.
Veröffentlicht: (2025)
Recurrent Diffusion for Large-Scale Parameter Generation
von: Wang, Kai, et al.
Veröffentlicht: (2025)
von: Wang, Kai, et al.
Veröffentlicht: (2025)
Steering LLMs via Scalable Interactive Oversight
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
SparseDM: Toward Sparse Efficient Diffusion Models
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
Efficient Controllable Diffusion via Optimal Classifier Guidance
von: Oertell, Owen, et al.
Veröffentlicht: (2025)
von: Oertell, Owen, et al.
Veröffentlicht: (2025)
Caracal: Causal Architecture via Spectral Mixing
von: Gan, Bingzheng, et al.
Veröffentlicht: (2026)
von: Gan, Bingzheng, et al.
Veröffentlicht: (2026)
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
von: Lin, Jinhong, et al.
Veröffentlicht: (2026)
von: Lin, Jinhong, et al.
Veröffentlicht: (2026)
Dynamically Scaled Activation Steering
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
Straight-Line Diffusion Model for Efficient 3D Molecular Generation
von: Ni, Yuyan, et al.
Veröffentlicht: (2025)
von: Ni, Yuyan, et al.
Veröffentlicht: (2025)
Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion Models
von: Dodson, Nick, et al.
Veröffentlicht: (2026)
von: Dodson, Nick, et al.
Veröffentlicht: (2026)
A Survey on Diffusion Models for Time Series and Spatio-Temporal Data
von: Yang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Yang, Yiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026) -
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
von: Kumar, Akash, et al.
Veröffentlicht: (2025) -
Out-of-Distribution Detection with a Single Unconditional Diffusion Model
von: Heng, Alvin, et al.
Veröffentlicht: (2024) -
DiffLoad: Uncertainty Quantification in Electrical Load Forecasting with the Diffusion Model
von: Wang, Zhixian, et al.
Veröffentlicht: (2023) -
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)