Saved in:
| Main Authors: | Rodriguez, Hector G., Rohrbach, Marcus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.25855 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
Chrono: A Simple Blueprint for Representing Time in MLLMs
by: Rodriguez, Hector, et al.
Published: (2024)
by: Rodriguez, Hector, et al.
Published: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
by: Arora, Aditya, et al.
Published: (2026)
by: Arora, Aditya, et al.
Published: (2026)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
by: Li, Zejun, et al.
Published: (2025)
by: Li, Zejun, et al.
Published: (2025)
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
by: Fan, Senran, et al.
Published: (2024)
by: Fan, Senran, et al.
Published: (2024)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
by: Han, Donghoon, et al.
Published: (2026)
by: Han, Donghoon, et al.
Published: (2026)
Distilling Specialized Orders for Visual Generation
by: Pramanik, Rishav, et al.
Published: (2025)
by: Pramanik, Rishav, et al.
Published: (2025)
Selective Visual Prompting in Vision Mamba
by: Yao, Yifeng, et al.
Published: (2024)
by: Yao, Yifeng, et al.
Published: (2024)
LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
by: Khare, Avishree, et al.
Published: (2025)
by: Khare, Avishree, et al.
Published: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
by: Georgenthum, Hugo, et al.
Published: (2025)
by: Georgenthum, Hugo, et al.
Published: (2025)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
by: Luo, Weiqing, et al.
Published: (2026)
by: Luo, Weiqing, et al.
Published: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
Visual Prompt Selection for In-Context Learning Segmentation
by: Suo, Wei, et al.
Published: (2024)
by: Suo, Wei, et al.
Published: (2024)
VideoScore2: Think before You Score in Generative Video Evaluation
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
Structured Visual Evidence Decomposition for Evidence-Grounded Multimodal Screening of Obstructive Sleep Apnea-Hypopnea Syndrome
by: Zhan, Chen, et al.
Published: (2026)
by: Zhan, Chen, et al.
Published: (2026)
DreamPolish: Domain Score Distillation With Progressive Geometry Generation
by: Cheng, Yean, et al.
Published: (2024)
by: Cheng, Yean, et al.
Published: (2024)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Text-to-3D Generation using Jensen-Shannon Score Distillation
by: Do, Khoi, et al.
Published: (2025)
by: Do, Khoi, et al.
Published: (2025)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
by: Li, Zeshang, et al.
Published: (2026)
by: Li, Zeshang, et al.
Published: (2026)
Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
by: Xie, Yutong, et al.
Published: (2026)
by: Xie, Yutong, et al.
Published: (2026)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
RS-Net: Context-Aware Relation Scoring for Dynamic Scene Graph Generation
by: Jo, Hae-Won, et al.
Published: (2025)
by: Jo, Hae-Won, et al.
Published: (2025)
Unconditional Human Motion and Shape Generation via Balanced Score-Based Diffusion
by: Björkstrand, David, et al.
Published: (2025)
by: Björkstrand, David, et al.
Published: (2025)
When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models
by: Chen, Qianpu, et al.
Published: (2026)
by: Chen, Qianpu, et al.
Published: (2026)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
by: Guo, Tengda, et al.
Published: (2026)
by: Guo, Tengda, et al.
Published: (2026)
Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding
by: Shao, Feixue, et al.
Published: (2026)
by: Shao, Feixue, et al.
Published: (2026)
Similar Items
-
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026) -
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023) -
Chrono: A Simple Blueprint for Representing Time in MLLMs
by: Rodriguez, Hector, et al.
Published: (2024) -
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)