Saved in:
| Main Authors: | D'Oronzio, Fabio, Putamorsi, Federico, Zini, Leonardo, Cornia, Marcella, Baraldi, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.25457 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
by: Baraldi, Lorenzo, et al.
Published: (2024)
by: Baraldi, Lorenzo, et al.
Published: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
by: Cocchi, Federico, et al.
Published: (2024)
by: Cocchi, Federico, et al.
Published: (2024)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
by: Turri, Evelyn, et al.
Published: (2026)
by: Turri, Evelyn, et al.
Published: (2026)
Tiny Inference-Time Scaling with Latent Verifiers
by: Bucciarelli, Davide, et al.
Published: (2026)
by: Bucciarelli, Davide, et al.
Published: (2026)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
by: Morini, Marco, et al.
Published: (2026)
by: Morini, Marco, et al.
Published: (2026)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
by: Cocchi, Federico, et al.
Published: (2025)
by: Cocchi, Federico, et al.
Published: (2025)
Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
by: Landi, Federico, et al.
Published: (2022)
by: Landi, Federico, et al.
Published: (2022)
Embodied Navigation at the Art Gallery
by: Bigazzi, Roberto, et al.
Published: (2022)
by: Bigazzi, Roberto, et al.
Published: (2022)
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
by: Poppi, Samuele, et al.
Published: (2023)
by: Poppi, Samuele, et al.
Published: (2023)
Explore and Explain: Self-supervised Navigation and Recounting
by: Bigazzi, Roberto, et al.
Published: (2020)
by: Bigazzi, Roberto, et al.
Published: (2020)
Multi-Class Unlearning for Image Classification via Weight Filtering
by: Poppi, Samuele, et al.
Published: (2023)
by: Poppi, Samuele, et al.
Published: (2023)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
by: Mattioli, Gabriele, et al.
Published: (2026)
by: Mattioli, Gabriele, et al.
Published: (2026)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
Recurrence Meets Transformers for Universal Multimodal Retrieval
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
by: Bucciarelli, Davide, et al.
Published: (2024)
by: Bucciarelli, Davide, et al.
Published: (2024)
The Revolution of Multimodal Large Language Models: A Survey
by: Caffagni, Davide, et al.
Published: (2024)
by: Caffagni, Davide, et al.
Published: (2024)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
by: Compagnoni, Alberto, et al.
Published: (2025)
by: Compagnoni, Alberto, et al.
Published: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
by: Caffagni, Davide, et al.
Published: (2024)
by: Caffagni, Davide, et al.
Published: (2024)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
by: Eddoubi, Hicham, et al.
Published: (2025)
by: Eddoubi, Hicham, et al.
Published: (2025)
Hallucination Early Detection in Diffusion Models
by: Betti, Federico, et al.
Published: (2026)
by: Betti, Federico, et al.
Published: (2026)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images
by: Amoroso, Roberto, et al.
Published: (2023)
by: Amoroso, Roberto, et al.
Published: (2023)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
by: Compagnoni, Alberto, et al.
Published: (2025)
by: Compagnoni, Alberto, et al.
Published: (2025)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
by: Betti, Federico, et al.
Published: (2024)
by: Betti, Federico, et al.
Published: (2024)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
SVGauge: Towards Human-Aligned Evaluation for SVG Generation
by: Zini, Leonardo, et al.
Published: (2025)
by: Zini, Leonardo, et al.
Published: (2025)
TextSR: Diffusion Super-Resolution with Multilingual OCR Guidance
by: Ye, Keren, et al.
Published: (2025)
by: Ye, Keren, et al.
Published: (2025)
TinySR: Pruning Diffusion for Real-World Image Super-Resolution
by: Dong, Linwei, et al.
Published: (2025)
by: Dong, Linwei, et al.
Published: (2025)
ICM-SR: Image-Conditioned Manifold Regularization for Image Super-Resolution
by: Kang, Junoh, et al.
Published: (2025)
by: Kang, Junoh, et al.
Published: (2025)
Similar Items
-
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
by: Baraldi, Lorenzo, et al.
Published: (2024) -
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023) -
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
by: Cocchi, Federico, et al.
Published: (2024) -
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
by: Sarto, Sara, et al.
Published: (2024) -
Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation
by: Barsellotti, Luca, et al.
Published: (2024)