SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Zhuohao, Jiang, Xingru, Gu, Weizheng, Wang, Yidong, Wen, Qingsong, Zhang, Shikun, Ye, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
RewardAnything: Generalizable Principle-Following Reward Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
di: Wang, Yidong, et al.
Pubblicazione: (2023)
di: Wang, Yidong, et al.
Pubblicazione: (2023)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
di: Zhou, Xiaoling, et al.
Pubblicazione: (2024)
di: Zhou, Xiaoling, et al.
Pubblicazione: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
Sparse Autoencoder Features for Classifications and Transferability
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
di: Muchane, Mark, et al.
Pubblicazione: (2025)
di: Muchane, Mark, et al.
Pubblicazione: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Steering Large Language Models for Machine Translation Personalization
di: Scalena, Daniel, et al.
Pubblicazione: (2025)
di: Scalena, Daniel, et al.
Pubblicazione: (2025)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
di: Mohamed, Asim, et al.
Pubblicazione: (2025)
di: Mohamed, Asim, et al.
Pubblicazione: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Team QUST at SemEval-2023 Task 3: A Comprehensive Study of Monolingual and Multilingual Approaches for Detecting Online News Genre, Framing and Persuasion Techniques
di: Jiang, Ye
Pubblicazione: (2023)
di: Jiang, Ye
Pubblicazione: (2023)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
di: Lv, Hang, et al.
Pubblicazione: (2025)
di: Lv, Hang, et al.
Pubblicazione: (2025)
Understanding and Mitigating Dataset Corruption in LLM Steering
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
Towards Understanding the Robustness of Sparse Autoencoders
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
di: Yu, Zhuohao, et al.
Pubblicazione: (2024) -
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026) -
RewardAnything: Generalizable Principle-Following Reward Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2025) -
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2024) -
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)