TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sanders, Kate, Weir, Nathaniel, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
von: Le, Van-Truong
Veröffentlicht: (2026)
von: Le, Van-Truong
Veröffentlicht: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
von: He, Wei
Veröffentlicht: (2026)
von: He, Wei
Veröffentlicht: (2026)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2022)
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2022)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
von: Dehghani, Mahshid, et al.
Veröffentlicht: (2024)
von: Dehghani, Mahshid, et al.
Veröffentlicht: (2024)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
von: Deichler, Anna, et al.
Veröffentlicht: (2026)
von: Deichler, Anna, et al.
Veröffentlicht: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
Defending against Backdoor Attacks via Module Switching
von: Li, Weijun, et al.
Veröffentlicht: (2025)
von: Li, Weijun, et al.
Veröffentlicht: (2025)
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
von: Teo, Charlton
Veröffentlicht: (2025)
von: Teo, Charlton
Veröffentlicht: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
Semantic Leakage from Image Embeddings
von: Chen, Yiyi, et al.
Veröffentlicht: (2026)
von: Chen, Yiyi, et al.
Veröffentlicht: (2026)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2021)
von: Parcalabescu, Letitia, et al.
Veröffentlicht: (2021)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
von: Deichler, Anna, et al.
Veröffentlicht: (2025)
von: Deichler, Anna, et al.
Veröffentlicht: (2025)
On the Limitations of Vision-Language Models in Understanding Image Transforms
von: Anis, Ahmad Mustafa, et al.
Veröffentlicht: (2025)
von: Anis, Ahmad Mustafa, et al.
Veröffentlicht: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
von: Singh, Sahajpreet, et al.
Veröffentlicht: (2025)
von: Singh, Sahajpreet, et al.
Veröffentlicht: (2025)
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
von: Sanders, Kate, et al.
Veröffentlicht: (2025)
von: Sanders, Kate, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025) -
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
von: Le, Van-Truong
Veröffentlicht: (2026) -
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026) -
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)