Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Franzmeyer, Tim, McAleer, Stephen, Henriques, João F., Foerster, Jakob N., Torr, Philip H. S., Bibi, Adel, de Witt, Christian Schroeder |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Select to Perfect: Imitating desired behavior from large multi-agent data
by: Franzmeyer, Tim, et al.
Published: (2024)
by: Franzmeyer, Tim, et al.
Published: (2024)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024)
by: Nasvytis, Linas, et al.
Published: (2024)
HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits
by: Franzmeyer, Tim, et al.
Published: (2024)
by: Franzmeyer, Tim, et al.
Published: (2024)
Towards Certification of Uncertainty Calibration under Adversarial Attacks
by: Emde, Cornelius, et al.
Published: (2024)
by: Emde, Cornelius, et al.
Published: (2024)
Plato's 'Republic'
by: McAleer, Sean
Published: (2020)
by: McAleer, Sean
Published: (2020)
TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
by: Nuti, Felipe, et al.
Published: (2025)
by: Nuti, Felipe, et al.
Published: (2025)
Faster Game Solving via Hyperparameter Schedules
by: Zhang, Naifeng, et al.
Published: (2024)
by: Zhang, Naifeng, et al.
Published: (2024)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
Algorithms and Complexity for Computing Nash Equilibria in Adversarial Team Games
by: Anagnostides, Ioannis, et al.
Published: (2023)
by: Anagnostides, Ioannis, et al.
Published: (2023)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
by: Lin, Runqi, et al.
Published: (2025)
by: Lin, Runqi, et al.
Published: (2025)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025)
by: Aljaafari, Tala, et al.
Published: (2025)
A Systematic Review to Explore Antenatal Care From the Perspectives of Women With Intellectual Disabilities and Midwives
by: Weam Alhulaibi, et al.
Published: (2024)
by: Weam Alhulaibi, et al.
Published: (2024)
Game-Theoretic Multiagent Reinforcement Learning
by: Yang, Yaodong, et al.
Published: (2020)
by: Yang, Yaodong, et al.
Published: (2020)
Prompting a Pretrained Transformer Can Be a Universal Approximator
by: Petrov, Aleksandar, et al.
Published: (2024)
by: Petrov, Aleksandar, et al.
Published: (2024)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
by: Petrov, Aleksandar, et al.
Published: (2023)
by: Petrov, Aleksandar, et al.
Published: (2023)
Tree Search for Language Model Agents
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
by: Kim, Hazel, et al.
Published: (2024)
by: Kim, Hazel, et al.
Published: (2024)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
by: Channing, Georgia, et al.
Published: (2024)
by: Channing, Georgia, et al.
Published: (2024)
Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
by: Liang, Yongyuan, et al.
Published: (2023)
by: Liang, Yongyuan, et al.
Published: (2023)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
by: Venhoff, Constantin, et al.
Published: (2024)
by: Venhoff, Constantin, et al.
Published: (2024)
Orienting and Engaging New Faculty
by: Pamela MacRae, et al.
Published: (2024)
by: Pamela MacRae, et al.
Published: (2024)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
The Danger Of Arrogance: Welfare Equilibra As A Solution To Stackelberg Self-Play In Non-Coincidental Games
by: Levi, Jake, et al.
Published: (2024)
by: Levi, Jake, et al.
Published: (2024)
Policy Space Response Oracles: A Survey
by: Bighashdel, Ariyan, et al.
Published: (2024)
by: Bighashdel, Ariyan, et al.
Published: (2024)
IPA-NeRF: Illusory Poisoning Attack Against Neural Radiance Fields
by: Jiang, Wenxiang, et al.
Published: (2024)
by: Jiang, Wenxiang, et al.
Published: (2024)
Select2Plan: Training-Free ICL-Based Planning through VQA and Memory Retrieval
by: Buoso, Davide, et al.
Published: (2024)
by: Buoso, Davide, et al.
Published: (2024)
A* Search Without Expansions: Learning Heuristic Functions with Deep Q-Networks
by: Agostinelli, Forest, et al.
Published: (2021)
by: Agostinelli, Forest, et al.
Published: (2021)
Sample-Efficient Regret-Minimizing Double Oracle in Extensive-Form Games
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection
by: Lakara, Kumud, et al.
Published: (2024)
by: Lakara, Kumud, et al.
Published: (2024)
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
by: Rose, Aaron, et al.
Published: (2026)
by: Rose, Aaron, et al.
Published: (2026)
Social Media Bot Policies: Evaluating Passive and Active Enforcement
by: Radivojevic, Kristina, et al.
Published: (2024)
by: Radivojevic, Kristina, et al.
Published: (2024)
Equivariant Networks for Zero-Shot Coordination
by: Muglich, Darius, et al.
Published: (2022)
by: Muglich, Darius, et al.
Published: (2022)
Ensemble Value Functions for Efficient Exploration in Multi-Agent Reinforcement Learning
by: Schäfer, Lukas, et al.
Published: (2023)
by: Schäfer, Lukas, et al.
Published: (2023)
On the Coexistence and Ensembling of Watermarks
by: Petrov, Aleksandar, et al.
Published: (2025)
by: Petrov, Aleksandar, et al.
Published: (2025)
Influencer Backdoor Attack on Semantic Segmentation
by: Lan, Haoheng, et al.
Published: (2023)
by: Lan, Haoheng, et al.
Published: (2023)
Computing Low-Entropy Couplings for Large-Support Distributions
by: Sokota, Samuel, et al.
Published: (2024)
by: Sokota, Samuel, et al.
Published: (2024)
Automated Design of Affine Maximizer Mechanisms in Dynamic Settings
by: Curry, Michael, et al.
Published: (2024)
by: Curry, Michael, et al.
Published: (2024)
Similar Items
-
Select to Perfect: Imitating desired behavior from large multi-agent data
by: Franzmeyer, Tim, et al.
Published: (2024) -
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024) -
HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits
by: Franzmeyer, Tim, et al.
Published: (2024) -
Towards Certification of Uncertainty Calibration under Adversarial Attacks
by: Emde, Cornelius, et al.
Published: (2024) -
Plato's 'Republic'
by: McAleer, Sean
Published: (2020)