Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Zhihao, Gong, Gracia, Zhu, Qinglin, Chen, Yudong, Zhao, Runcong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
von: Zhu, Qinglin, et al.
Veröffentlicht: (2025)
von: Zhu, Qinglin, et al.
Veröffentlicht: (2025)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
von: Qi, Siya, et al.
Veröffentlicht: (2026)
von: Qi, Siya, et al.
Veröffentlicht: (2026)
PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
von: Zhu, Qinglin, et al.
Veröffentlicht: (2024)
von: Zhu, Qinglin, et al.
Veröffentlicht: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
von: Hu, Zhanghao, et al.
Veröffentlicht: (2026)
von: Hu, Zhanghao, et al.
Veröffentlicht: (2026)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
von: Zhao, Runcong, et al.
Veröffentlicht: (2024)
von: Zhao, Runcong, et al.
Veröffentlicht: (2024)
Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding
von: Zhu, Lixing, et al.
Veröffentlicht: (2023)
von: Zhu, Lixing, et al.
Veröffentlicht: (2023)
Large Scale Knowledge Washing
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Sparse Activation Editing for Reliable Instruction Following in Narratives
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
von: Zhu, Qinglin, et al.
Veröffentlicht: (2025)
von: Zhu, Qinglin, et al.
Veröffentlicht: (2025)
More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles
von: Chen, Ruibo, et al.
Veröffentlicht: (2026)
von: Chen, Ruibo, et al.
Veröffentlicht: (2026)
Ensemble Watermarks for Large Language Models
von: Niess, Georg, et al.
Veröffentlicht: (2024)
von: Niess, Georg, et al.
Veröffentlicht: (2024)
Pull Requests as a Training Signal for Repo-Level Code Editing
von: Zhu, Qinglin, et al.
Veröffentlicht: (2026)
von: Zhu, Qinglin, et al.
Veröffentlicht: (2026)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
Few-shot LLM Synthetic Data with Distribution Matching
von: Ren, Jiyuan, et al.
Veröffentlicht: (2025)
von: Ren, Jiyuan, et al.
Veröffentlicht: (2025)
Fragile Reasoning: A Mechanistic Analysis of LLM Sensitivity to Meaning-Preserving Perturbations
von: Han, Shou-Tzu, et al.
Veröffentlicht: (2026)
von: Han, Shou-Tzu, et al.
Veröffentlicht: (2026)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMs
von: Zhu, Junlin, et al.
Veröffentlicht: (2026)
von: Zhu, Junlin, et al.
Veröffentlicht: (2026)
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment
von: Li, Jiazheng, et al.
Veröffentlicht: (2024)
von: Li, Jiazheng, et al.
Veröffentlicht: (2024)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
Towards Codable Watermarking for Injecting Multi-bits Information to LLMs
von: Wang, Lean, et al.
Veröffentlicht: (2023)
von: Wang, Lean, et al.
Veröffentlicht: (2023)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
von: Zhuo, Anyuan, et al.
Veröffentlicht: (2025)
von: Zhuo, Anyuan, et al.
Veröffentlicht: (2025)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
Don't Throw Away Your Pretrained Model
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
von: Xu, Yixuan, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan, et al.
Veröffentlicht: (2025)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
SoK: Are Watermarks in LLMs Ready for Deployment?
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Improved Unbiased Watermark for Large Language Models
von: Chen, Ruibo, et al.
Veröffentlicht: (2025)
von: Chen, Ruibo, et al.
Veröffentlicht: (2025)
A Watermark for Order-Agnostic Language Models
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
von: Fadeeva, Ekaterina, et al.
Veröffentlicht: (2025)
von: Fadeeva, Ekaterina, et al.
Veröffentlicht: (2025)
Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs
von: Luo, Yiyang, et al.
Veröffentlicht: (2024)
von: Luo, Yiyang, et al.
Veröffentlicht: (2024)
De-mark: Watermark Removal in Large Language Models
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
von: Hoang, Duy C., et al.
Veröffentlicht: (2024)
von: Hoang, Duy C., et al.
Veröffentlicht: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
von: Zhu, Qinglin, et al.
Veröffentlicht: (2025) -
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
von: Qi, Siya, et al.
Veröffentlicht: (2026) -
PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
von: Zhu, Qinglin, et al.
Veröffentlicht: (2024) -
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
von: Zhao, Runcong, et al.
Veröffentlicht: (2025) -
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)