Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
Fuente:
arXiv
Guardado en:
| Autores principales: | Masala, Mihai, Leordeanu, Marius |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
por: Masala, Mihai, et al.
Publicado: (2025)
por: Masala, Mihai, et al.
Publicado: (2025)
Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
por: Cudlenco, Nicolae, et al.
Publicado: (2026)
por: Cudlenco, Nicolae, et al.
Publicado: (2026)
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
por: Cudlenco, Nicolae, et al.
Publicado: (2026)
por: Cudlenco, Nicolae, et al.
Publicado: (2026)
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
por: Airinei, Daniel, et al.
Publicado: (2025)
por: Airinei, Daniel, et al.
Publicado: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
Fostering Video Reasoning via Next-Event Prediction
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
por: Jiang, Ruixiang, et al.
Publicado: (2025)
por: Jiang, Ruixiang, et al.
Publicado: (2025)
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
por: Mihai-Cristian, Pîrvu, et al.
Publicado: (2025)
por: Mihai-Cristian, Pîrvu, et al.
Publicado: (2025)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
por: Gu, Zishan, et al.
Publicado: (2024)
por: Gu, Zishan, et al.
Publicado: (2024)
Commonsense for Zero-Shot Natural Language Video Localization
por: Holla, Meghana, et al.
Publicado: (2023)
por: Holla, Meghana, et al.
Publicado: (2023)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
por: Taguchi, Shun, et al.
Publicado: (2025)
por: Taguchi, Shun, et al.
Publicado: (2025)
DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
por: Hong, Susung, et al.
Publicado: (2023)
por: Hong, Susung, et al.
Publicado: (2023)
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
por: Costea, Dragos, et al.
Publicado: (2026)
por: Costea, Dragos, et al.
Publicado: (2026)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
por: Nagar, Aishik, et al.
Publicado: (2024)
por: Nagar, Aishik, et al.
Publicado: (2024)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
por: Yang, Bang, et al.
Publicado: (2023)
por: Yang, Bang, et al.
Publicado: (2023)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing
por: Jeong, Hyeonho, et al.
Publicado: (2024)
por: Jeong, Hyeonho, et al.
Publicado: (2024)
Visually Descriptive Language Model for Vector Graphics Reasoning
por: Wang, Zhenhailong, et al.
Publicado: (2024)
por: Wang, Zhenhailong, et al.
Publicado: (2024)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
por: Chen, Jiali, et al.
Publicado: (2024)
por: Chen, Jiali, et al.
Publicado: (2024)
CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
por: Sarkar, Sayan Deb, et al.
Publicado: (2026)
por: Sarkar, Sayan Deb, et al.
Publicado: (2026)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
por: Chen, Chao, et al.
Publicado: (2025)
por: Chen, Chao, et al.
Publicado: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
por: Liang, Xiwen, et al.
Publicado: (2023)
por: Liang, Xiwen, et al.
Publicado: (2023)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
por: Wang, Ziyang, et al.
Publicado: (2024)
por: Wang, Ziyang, et al.
Publicado: (2024)
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
por: Agarwal, Lakshita, et al.
Publicado: (2025)
por: Agarwal, Lakshita, et al.
Publicado: (2025)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
por: Byun, Ji Young, et al.
Publicado: (2025)
por: Byun, Ji Young, et al.
Publicado: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
por: Wang, Haozhe, et al.
Publicado: (2025)
por: Wang, Haozhe, et al.
Publicado: (2025)
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
por: Song, Xiujie, et al.
Publicado: (2024)
por: Song, Xiujie, et al.
Publicado: (2024)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
por: Qu, Kevin, et al.
Publicado: (2026)
por: Qu, Kevin, et al.
Publicado: (2026)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
por: Allgeuer, Philipp, et al.
Publicado: (2024)
por: Allgeuer, Philipp, et al.
Publicado: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
por: Zhang, Jianrui, et al.
Publicado: (2024)
por: Zhang, Jianrui, et al.
Publicado: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
por: Feng, Sicheng, et al.
Publicado: (2025)
por: Feng, Sicheng, et al.
Publicado: (2025)
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
por: Zhang, Fanrui, et al.
Publicado: (2025)
por: Zhang, Fanrui, et al.
Publicado: (2025)
TimeRefine: Temporal Grounding with Time Refining Video LLM
por: Wang, Xizi, et al.
Publicado: (2024)
por: Wang, Xizi, et al.
Publicado: (2024)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
por: Ananthram, Amith, et al.
Publicado: (2025)
por: Ananthram, Amith, et al.
Publicado: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
Ejemplares similares
-
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
por: Masala, Mihai, et al.
Publicado: (2025) -
Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
por: Cudlenco, Nicolae, et al.
Publicado: (2026) -
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
por: Cudlenco, Nicolae, et al.
Publicado: (2026) -
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
por: Airinei, Daniel, et al.
Publicado: (2025) -
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
por: Guo, Ziyu, et al.
Publicado: (2025)