Guardado en:
| Autores principales: | Farzam, Amirhossein, Behabahani, Majid, Malek, Mani, Nevmyvaka, Yuriy, Sapiro, Guillermo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.19396 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs
por: Reza, Md Farhamdur, et al.
Publicado: (2026)
por: Reza, Md Farhamdur, et al.
Publicado: (2026)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
por: Geng, Jianing, et al.
Publicado: (2025)
por: Geng, Jianing, et al.
Publicado: (2025)
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
por: Schlesinger, Oded, et al.
Publicado: (2025)
por: Schlesinger, Oded, et al.
Publicado: (2025)
Antelope: Potent and Concealed Jailbreak Attack Strategy
por: Zhao, Xin, et al.
Publicado: (2024)
por: Zhao, Xin, et al.
Publicado: (2024)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
por: Cui, Tiehan, et al.
Publicado: (2025)
por: Cui, Tiehan, et al.
Publicado: (2025)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
por: Hartman, Max, et al.
Publicado: (2026)
por: Hartman, Max, et al.
Publicado: (2026)
Data-Aware Random Feature Kernel for Transformers
por: Farzam, Amirhossein, et al.
Publicado: (2026)
por: Farzam, Amirhossein, et al.
Publicado: (2026)
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
por: Vaidhya, Tejas, et al.
Publicado: (2025)
por: Vaidhya, Tejas, et al.
Publicado: (2025)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
por: Roger, Alexis, et al.
Publicado: (2025)
por: Roger, Alexis, et al.
Publicado: (2025)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
por: Timmer, Roelien C., et al.
Publicado: (2024)
por: Timmer, Roelien C., et al.
Publicado: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
por: Chen, Depeng, et al.
Publicado: (2024)
por: Chen, Depeng, et al.
Publicado: (2024)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
por: Wang, Zhilong, et al.
Publicado: (2024)
por: Wang, Zhilong, et al.
Publicado: (2024)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
por: Hogan, Brendan R., et al.
Publicado: (2026)
por: Hogan, Brendan R., et al.
Publicado: (2026)
Wrist Photoplethysmography Predicts Dietary Information
por: Verrier, Kyle, et al.
Publicado: (2025)
por: Verrier, Kyle, et al.
Publicado: (2025)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
por: Garg, Sahil, et al.
Publicado: (2024)
por: Garg, Sahil, et al.
Publicado: (2024)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
por: Riachi, Roland, et al.
Publicado: (2025)
por: Riachi, Roland, et al.
Publicado: (2025)
Federated Fairness without Access to Sensitive Groups
por: Papadaki, Afroditi, et al.
Publicado: (2024)
por: Papadaki, Afroditi, et al.
Publicado: (2024)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
por: Park, Yein, et al.
Publicado: (2025)
por: Park, Yein, et al.
Publicado: (2025)
Hide to Guide: Learning via Semantic Masking
por: Liu, Ruitao, et al.
Publicado: (2026)
por: Liu, Ruitao, et al.
Publicado: (2026)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
por: Alanova, Shirin, et al.
Publicado: (2026)
por: Alanova, Shirin, et al.
Publicado: (2026)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
por: Lee, Wonjun, et al.
Publicado: (2025)
por: Lee, Wonjun, et al.
Publicado: (2025)
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
por: Ning, Kanghui, et al.
Publicado: (2025)
por: Ning, Kanghui, et al.
Publicado: (2025)
Activation-Guided Local Editing for Jailbreaking Attacks
por: Wang, Jiecong, et al.
Publicado: (2025)
por: Wang, Jiecong, et al.
Publicado: (2025)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
por: Chen, Yanxi, et al.
Publicado: (2025)
por: Chen, Yanxi, et al.
Publicado: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
por: Ji, Haoxuan, et al.
Publicado: (2024)
por: Ji, Haoxuan, et al.
Publicado: (2024)
Graph Partitioning With Limited Moves
por: Behbahani, Majid, et al.
Publicado: (2024)
por: Behbahani, Majid, et al.
Publicado: (2024)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
por: Guo, Weiyang, et al.
Publicado: (2025)
por: Guo, Weiyang, et al.
Publicado: (2025)
Subgoal Discovery Using a Free Energy Paradigm and State Aggregations
por: Mesbah, Amirhossein, et al.
Publicado: (2024)
por: Mesbah, Amirhossein, et al.
Publicado: (2024)
Jailbreaking to Jailbreak
por: Kritz, Jeremy, et al.
Publicado: (2025)
por: Kritz, Jeremy, et al.
Publicado: (2025)
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
por: Kim, Minseon, et al.
Publicado: (2024)
por: Kim, Minseon, et al.
Publicado: (2024)
Plan-X: Instruct Video Generation via Semantic Planning
por: Huang, Lun, et al.
Publicado: (2025)
por: Huang, Lun, et al.
Publicado: (2025)
Concealed Adversarial attacks on neural networks for sequential data
por: Sokerin, Petr, et al.
Publicado: (2025)
por: Sokerin, Petr, et al.
Publicado: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
por: Kalibhat, Neha, et al.
Publicado: (2026)
por: Kalibhat, Neha, et al.
Publicado: (2026)
An Empirical Evaluation of Neural and Neuro-symbolic Approaches to Real-time Multimodal Complex Event Detection
por: Han, Liying, et al.
Publicado: (2024)
por: Han, Liying, et al.
Publicado: (2024)
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images
por: Daw, Arka, et al.
Publicado: (2024)
por: Daw, Arka, et al.
Publicado: (2024)
Segment Concealed Objects with Incomplete Supervision
por: He, Chunming, et al.
Publicado: (2025)
por: He, Chunming, et al.
Publicado: (2025)
Interpretable Discriminative Text Representations via Agreement and Label Disentanglement
por: Wang, Tong, et al.
Publicado: (2026)
por: Wang, Tong, et al.
Publicado: (2026)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
por: Ashok, Dhananjay, et al.
Publicado: (2026)
por: Ashok, Dhananjay, et al.
Publicado: (2026)
Sycophancy Hides Linearly in the Attention Heads
por: Genadi, Rifo, et al.
Publicado: (2026)
por: Genadi, Rifo, et al.
Publicado: (2026)
Ejemplares similares
-
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs
por: Reza, Md Farhamdur, et al.
Publicado: (2026) -
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
por: Geng, Jianing, et al.
Publicado: (2025) -
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
por: Schlesinger, Oded, et al.
Publicado: (2025) -
Antelope: Potent and Concealed Jailbreak Attack Strategy
por: Zhao, Xin, et al.
Publicado: (2024) -
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
por: Cui, Tiehan, et al.
Publicado: (2025)