Salvato in:
| Autori principali: | Arora, Aditya, Gupta, Akshita, Rodriguez, Pau, Rohrbach, Marcus |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.18575 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReCap: Better Gaussian Relighting with Cross-Environment Captures
di: Li, Jingzhi, et al.
Pubblicazione: (2024)
di: Li, Jingzhi, et al.
Pubblicazione: (2024)
Video ReCap: Recursive Captioning of Hour-Long Videos
di: Islam, Md Mohaiminul, et al.
Pubblicazione: (2024)
di: Islam, Md Mohaiminul, et al.
Pubblicazione: (2024)
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
di: Nguyen, Thinh-Phuc, et al.
Pubblicazione: (2025)
di: Nguyen, Thinh-Phuc, et al.
Pubblicazione: (2025)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2026)
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2026)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
di: Rodriguez, Hector G., et al.
Pubblicazione: (2026)
di: Rodriguez, Hector G., et al.
Pubblicazione: (2026)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
di: Abdessaied, Adnen, et al.
Pubblicazione: (2025)
di: Abdessaied, Adnen, et al.
Pubblicazione: (2025)
Chrono: A Simple Blueprint for Representing Time in MLLMs
di: Rodriguez, Hector, et al.
Pubblicazione: (2024)
di: Rodriguez, Hector, et al.
Pubblicazione: (2024)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
di: Braun, Tobias, et al.
Pubblicazione: (2024)
di: Braun, Tobias, et al.
Pubblicazione: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
di: Wang, Yabing, et al.
Pubblicazione: (2024)
di: Wang, Yabing, et al.
Pubblicazione: (2024)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
di: Tao, Ming, et al.
Pubblicazione: (2024)
di: Tao, Ming, et al.
Pubblicazione: (2024)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
di: Sarkar, Ayushman, et al.
Pubblicazione: (2026)
di: Sarkar, Ayushman, et al.
Pubblicazione: (2026)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
Grounding Language in Multi-Perspective Referential Communication
di: Tang, Zineng, et al.
Pubblicazione: (2024)
di: Tang, Zineng, et al.
Pubblicazione: (2024)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2025)
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2025)
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
di: Ma, Tianren, et al.
Pubblicazione: (2024)
di: Ma, Tianren, et al.
Pubblicazione: (2024)
Predicting Implicit Arguments in Procedural Video Instructions
di: Batra, Anil, et al.
Pubblicazione: (2025)
di: Batra, Anil, et al.
Pubblicazione: (2025)
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
di: Sun, Ye, et al.
Pubblicazione: (2025)
di: Sun, Ye, et al.
Pubblicazione: (2025)
SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks
di: Wang, Shining, et al.
Pubblicazione: (2025)
di: Wang, Shining, et al.
Pubblicazione: (2025)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
di: Deng, Tianchen, et al.
Pubblicazione: (2025)
di: Deng, Tianchen, et al.
Pubblicazione: (2025)
"Previously on ..." From Recaps to Story Summarization
di: Singh, Aditya Kumar, et al.
Pubblicazione: (2024)
di: Singh, Aditya Kumar, et al.
Pubblicazione: (2024)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
di: Saxena, Pranav, et al.
Pubblicazione: (2025)
di: Saxena, Pranav, et al.
Pubblicazione: (2025)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
di: Li, Binbin, et al.
Pubblicazione: (2025)
di: Li, Binbin, et al.
Pubblicazione: (2025)
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
di: Zhuang, Cailin, et al.
Pubblicazione: (2025)
di: Zhuang, Cailin, et al.
Pubblicazione: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
di: Shen, Xiaoqian, et al.
Pubblicazione: (2023)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2023)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
di: Narnaware, Vishal, et al.
Pubblicazione: (2026)
di: Narnaware, Vishal, et al.
Pubblicazione: (2026)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
di: An, Zhaochong, et al.
Pubblicazione: (2025)
di: An, Zhaochong, et al.
Pubblicazione: (2025)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
di: Mao, Jiawei, et al.
Pubblicazione: (2024)
di: Mao, Jiawei, et al.
Pubblicazione: (2024)
ReGround: Improving Textual and Spatial Grounding at No Cost
di: Lee, Phillip Y., et al.
Pubblicazione: (2024)
di: Lee, Phillip Y., et al.
Pubblicazione: (2024)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
di: T V, Sethuraman, et al.
Pubblicazione: (2026)
di: T V, Sethuraman, et al.
Pubblicazione: (2026)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
di: Jeong, Yujin, et al.
Pubblicazione: (2025)
di: Jeong, Yujin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ReCap: Better Gaussian Relighting with Cross-Environment Captures
di: Li, Jingzhi, et al.
Pubblicazione: (2024) -
Video ReCap: Recursive Captioning of Hour-Long Videos
di: Islam, Md Mohaiminul, et al.
Pubblicazione: (2024) -
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
di: Nguyen, Thinh-Phuc, et al.
Pubblicazione: (2025) -
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2026) -
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
di: Rodriguez, Hector G., et al.
Pubblicazione: (2026)