Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Zhihao, Gong, Gracia, Zhu, Qinglin, Chen, Yudong, Zhao, Runcong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
por: Zhu, Qinglin, et al.
Publicado: (2025)
por: Zhu, Qinglin, et al.
Publicado: (2025)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
por: Qi, Siya, et al.
Publicado: (2026)
por: Qi, Siya, et al.
Publicado: (2026)
PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
por: Zhu, Qinglin, et al.
Publicado: (2024)
por: Zhu, Qinglin, et al.
Publicado: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
por: Zhao, Runcong, et al.
Publicado: (2025)
por: Zhao, Runcong, et al.
Publicado: (2025)
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
por: Potraghloo, Erfan Baghaei, et al.
Publicado: (2026)
por: Potraghloo, Erfan Baghaei, et al.
Publicado: (2026)
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
por: Hu, Zhanghao, et al.
Publicado: (2026)
por: Hu, Zhanghao, et al.
Publicado: (2026)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
por: Zhao, Runcong, et al.
Publicado: (2024)
por: Zhao, Runcong, et al.
Publicado: (2024)
Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding
por: Zhu, Lixing, et al.
Publicado: (2023)
por: Zhu, Lixing, et al.
Publicado: (2023)
Large Scale Knowledge Washing
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
Sparse Activation Editing for Reliable Instruction Following in Narratives
por: Zhao, Runcong, et al.
Publicado: (2025)
por: Zhao, Runcong, et al.
Publicado: (2025)
Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
por: Zhu, Qinglin, et al.
Publicado: (2025)
por: Zhu, Qinglin, et al.
Publicado: (2025)
More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles
por: Chen, Ruibo, et al.
Publicado: (2026)
por: Chen, Ruibo, et al.
Publicado: (2026)
Ensemble Watermarks for Large Language Models
por: Niess, Georg, et al.
Publicado: (2024)
por: Niess, Georg, et al.
Publicado: (2024)
Pull Requests as a Training Signal for Repo-Level Code Editing
por: Zhu, Qinglin, et al.
Publicado: (2026)
por: Zhu, Qinglin, et al.
Publicado: (2026)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
por: Xu, Hainiu, et al.
Publicado: (2024)
por: Xu, Hainiu, et al.
Publicado: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
por: Wu, Xinwei, et al.
Publicado: (2025)
por: Wu, Xinwei, et al.
Publicado: (2025)
Few-shot LLM Synthetic Data with Distribution Matching
por: Ren, Jiyuan, et al.
Publicado: (2025)
por: Ren, Jiyuan, et al.
Publicado: (2025)
Fragile Reasoning: A Mechanistic Analysis of LLM Sensitivity to Meaning-Preserving Perturbations
por: Han, Shou-Tzu, et al.
Publicado: (2026)
por: Han, Shou-Tzu, et al.
Publicado: (2026)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
por: Sternfeld, Alexander, et al.
Publicado: (2026)
por: Sternfeld, Alexander, et al.
Publicado: (2026)
QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMs
por: Zhu, Junlin, et al.
Publicado: (2026)
por: Zhu, Junlin, et al.
Publicado: (2026)
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment
por: Li, Jiazheng, et al.
Publicado: (2024)
por: Li, Jiazheng, et al.
Publicado: (2024)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
por: Shetty, Pranav, et al.
Publicado: (2025)
por: Shetty, Pranav, et al.
Publicado: (2025)
Towards Codable Watermarking for Injecting Multi-bits Information to LLMs
por: Wang, Lean, et al.
Publicado: (2023)
por: Wang, Lean, et al.
Publicado: (2023)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
por: Liu, Yepeng, et al.
Publicado: (2025)
por: Liu, Yepeng, et al.
Publicado: (2025)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
por: Zhuo, Anyuan, et al.
Publicado: (2025)
por: Zhuo, Anyuan, et al.
Publicado: (2025)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
por: Zhao, Runcong, et al.
Publicado: (2025)
por: Zhao, Runcong, et al.
Publicado: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
por: Kassem, Aly M., et al.
Publicado: (2025)
por: Kassem, Aly M., et al.
Publicado: (2025)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
Don't Throw Away Your Pretrained Model
por: Feng, Shangbin, et al.
Publicado: (2025)
por: Feng, Shangbin, et al.
Publicado: (2025)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
por: Xu, Yixuan, et al.
Publicado: (2025)
por: Xu, Yixuan, et al.
Publicado: (2025)
Watermarking LLM Agent Trajectories
por: Meng, Wenlong, et al.
Publicado: (2026)
por: Meng, Wenlong, et al.
Publicado: (2026)
SoK: Are Watermarks in LLMs Ready for Deployment?
por: Dang, Kieu, et al.
Publicado: (2025)
por: Dang, Kieu, et al.
Publicado: (2025)
Improved Unbiased Watermark for Large Language Models
por: Chen, Ruibo, et al.
Publicado: (2025)
por: Chen, Ruibo, et al.
Publicado: (2025)
A Watermark for Order-Agnostic Language Models
por: Chen, Ruibo, et al.
Publicado: (2024)
por: Chen, Ruibo, et al.
Publicado: (2024)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs
por: Luo, Yiyang, et al.
Publicado: (2024)
por: Luo, Yiyang, et al.
Publicado: (2024)
De-mark: Watermark Removal in Large Language Models
por: Chen, Ruibo, et al.
Publicado: (2024)
por: Chen, Ruibo, et al.
Publicado: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
por: Poppi, Samuele, et al.
Publicado: (2024)
por: Poppi, Samuele, et al.
Publicado: (2024)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
por: Hoang, Duy C., et al.
Publicado: (2024)
por: Hoang, Duy C., et al.
Publicado: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
por: Wang, Jun, et al.
Publicado: (2024)
por: Wang, Jun, et al.
Publicado: (2024)
Ejemplares similares
-
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
por: Zhu, Qinglin, et al.
Publicado: (2025) -
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
por: Qi, Siya, et al.
Publicado: (2026) -
PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
por: Zhu, Qinglin, et al.
Publicado: (2024) -
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
por: Zhao, Runcong, et al.
Publicado: (2025) -
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
por: Potraghloo, Erfan Baghaei, et al.
Publicado: (2026)