RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Eddoubi, Hicham, Ricker, Jonas, Cocchi, Federico, Baraldi, Lorenzo, Sotgiu, Angelo, Pintor, Maura, Cornia, Marcella, Fischer, Asja, Cucchiara, Rita, Biggio, Battista |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
Tiny Inference-Time Scaling with Latent Verifiers
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
The Revolution of Multimodal Large Language Models: A Survey
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Embodied Agents for Efficient Exploration and Smart Scene Description
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2023)
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2023)
Recurrence Meets Transformers for Universal Multimodal Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
Multi-Class Unlearning for Image Classification via Weight Filtering
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
von: Landi, Federico, et al.
Veröffentlicht: (2022)
von: Landi, Federico, et al.
Veröffentlicht: (2022)
Explore and Explain: Self-supervised Navigation and Recounting
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2020)
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2020)
Embodied Navigation at the Art Gallery
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2022)
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2022)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
von: Betti, Federico, et al.
Veröffentlicht: (2024)
von: Betti, Federico, et al.
Veröffentlicht: (2024)
Hallucination Early Detection in Diffusion Models
von: Betti, Federico, et al.
Veröffentlicht: (2026)
von: Betti, Federico, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
von: Pintor, Maura, et al.
Veröffentlicht: (2022)
von: Pintor, Maura, et al.
Veröffentlicht: (2022)
Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images
von: Amoroso, Roberto, et al.
Veröffentlicht: (2023)
von: Amoroso, Roberto, et al.
Veröffentlicht: (2023)
AIGeN: An Adversarial Approach for Instruction Generation in VLN
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
von: Morini, Marco, et al.
Veröffentlicht: (2026)
von: Morini, Marco, et al.
Veröffentlicht: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution
von: D'Oronzio, Fabio, et al.
Veröffentlicht: (2026)
von: D'Oronzio, Fabio, et al.
Veröffentlicht: (2026)
UNMuTe: Unifying Navigation and Multimodal Dialogue-like Text Generation
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
von: Turri, Evelyn, et al.
Veröffentlicht: (2026)
von: Turri, Evelyn, et al.
Veröffentlicht: (2026)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
von: Cappelletti, Silvia, et al.
Veröffentlicht: (2025)
von: Cappelletti, Silvia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2024) -
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024) -
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
von: Poppi, Samuele, et al.
Veröffentlicht: (2023) -
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024) -
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
von: Sarto, Sara, et al.
Veröffentlicht: (2024)