Gespeichert in:
| Hauptverfasser: | Farzam, Amirhossein, Behabahani, Majid, Malek, Mani, Nevmyvaka, Yuriy, Sapiro, Guillermo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.19396 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs
von: Reza, Md Farhamdur, et al.
Veröffentlicht: (2026)
von: Reza, Md Farhamdur, et al.
Veröffentlicht: (2026)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
von: Schlesinger, Oded, et al.
Veröffentlicht: (2025)
von: Schlesinger, Oded, et al.
Veröffentlicht: (2025)
Antelope: Potent and Concealed Jailbreak Attack Strategy
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
von: Cui, Tiehan, et al.
Veröffentlicht: (2025)
von: Cui, Tiehan, et al.
Veröffentlicht: (2025)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
von: Hartman, Max, et al.
Veröffentlicht: (2026)
von: Hartman, Max, et al.
Veröffentlicht: (2026)
Data-Aware Random Feature Kernel for Transformers
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
von: Vaidhya, Tejas, et al.
Veröffentlicht: (2025)
von: Vaidhya, Tejas, et al.
Veröffentlicht: (2025)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
von: Timmer, Roelien C., et al.
Veröffentlicht: (2024)
von: Timmer, Roelien C., et al.
Veröffentlicht: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
von: Wang, Zhilong, et al.
Veröffentlicht: (2024)
von: Wang, Zhilong, et al.
Veröffentlicht: (2024)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
von: Hogan, Brendan R., et al.
Veröffentlicht: (2026)
von: Hogan, Brendan R., et al.
Veröffentlicht: (2026)
Wrist Photoplethysmography Predicts Dietary Information
von: Verrier, Kyle, et al.
Veröffentlicht: (2025)
von: Verrier, Kyle, et al.
Veröffentlicht: (2025)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
von: Garg, Sahil, et al.
Veröffentlicht: (2024)
von: Garg, Sahil, et al.
Veröffentlicht: (2024)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
Federated Fairness without Access to Sensitive Groups
von: Papadaki, Afroditi, et al.
Veröffentlicht: (2024)
von: Papadaki, Afroditi, et al.
Veröffentlicht: (2024)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
Hide to Guide: Learning via Semantic Masking
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
von: Alanova, Shirin, et al.
Veröffentlicht: (2026)
von: Alanova, Shirin, et al.
Veröffentlicht: (2026)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
von: Ning, Kanghui, et al.
Veröffentlicht: (2025)
von: Ning, Kanghui, et al.
Veröffentlicht: (2025)
Activation-Guided Local Editing for Jailbreaking Attacks
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
Graph Partitioning With Limited Moves
von: Behbahani, Majid, et al.
Veröffentlicht: (2024)
von: Behbahani, Majid, et al.
Veröffentlicht: (2024)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
Subgoal Discovery Using a Free Energy Paradigm and State Aggregations
von: Mesbah, Amirhossein, et al.
Veröffentlicht: (2024)
von: Mesbah, Amirhossein, et al.
Veröffentlicht: (2024)
Jailbreaking to Jailbreak
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
von: Kim, Minseon, et al.
Veröffentlicht: (2024)
von: Kim, Minseon, et al.
Veröffentlicht: (2024)
Plan-X: Instruct Video Generation via Semantic Planning
von: Huang, Lun, et al.
Veröffentlicht: (2025)
von: Huang, Lun, et al.
Veröffentlicht: (2025)
Concealed Adversarial attacks on neural networks for sequential data
von: Sokerin, Petr, et al.
Veröffentlicht: (2025)
von: Sokerin, Petr, et al.
Veröffentlicht: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of Neural and Neuro-symbolic Approaches to Real-time Multimodal Complex Event Detection
von: Han, Liying, et al.
Veröffentlicht: (2024)
von: Han, Liying, et al.
Veröffentlicht: (2024)
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images
von: Daw, Arka, et al.
Veröffentlicht: (2024)
von: Daw, Arka, et al.
Veröffentlicht: (2024)
Segment Concealed Objects with Incomplete Supervision
von: He, Chunming, et al.
Veröffentlicht: (2025)
von: He, Chunming, et al.
Veröffentlicht: (2025)
Interpretable Discriminative Text Representations via Agreement and Label Disentanglement
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs
von: Reza, Md Farhamdur, et al.
Veröffentlicht: (2026) -
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025) -
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
von: Schlesinger, Oded, et al.
Veröffentlicht: (2025) -
Antelope: Potent and Concealed Jailbreak Attack Strategy
von: Zhao, Xin, et al.
Veröffentlicht: (2024) -
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
von: Cui, Tiehan, et al.
Veröffentlicht: (2025)