RECODE: Reasoning Through Code Generation for Visual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Junhong, Cai, Mu, Hu, Bo, Talwalkar, Ameet, Ross, David A, Schmid, Cordelia, Fathi, Alireza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
von: Xie, Stephan, et al.
Veröffentlicht: (2026)
von: Xie, Stephan, et al.
Veröffentlicht: (2026)
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
von: Min, Juhong, et al.
Veröffentlicht: (2024)
von: Min, Juhong, et al.
Veröffentlicht: (2024)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
von: Huang, Ian, et al.
Veröffentlicht: (2025)
von: Huang, Ian, et al.
Veröffentlicht: (2025)
Language-Guided Image Tokenization for Generation
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
von: Hu, Ziniu, et al.
Veröffentlicht: (2024)
von: Hu, Ziniu, et al.
Veröffentlicht: (2024)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
Specialized Foundation Models Struggle to Beat Supervised Baselines
von: Xu, Zongzhe, et al.
Veröffentlicht: (2024)
von: Xu, Zongzhe, et al.
Veröffentlicht: (2024)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
von: Nguyen, Jason, et al.
Veröffentlicht: (2026)
von: Nguyen, Jason, et al.
Veröffentlicht: (2026)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
von: Lassoued, Aymen, et al.
Veröffentlicht: (2026)
von: Lassoued, Aymen, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
Geospatial Chain of Thought Reasoning for Enhanced Visual Question Answering on Satellite Imagery
von: Shanker, Shambhavi, et al.
Veröffentlicht: (2025)
von: Shanker, Shambhavi, et al.
Veröffentlicht: (2025)
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
von: Jing, Liu, et al.
Veröffentlicht: (2025)
von: Jing, Liu, et al.
Veröffentlicht: (2025)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
Questioning the Stability of Visual Question Answering
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024) -
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024) -
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
von: Xie, Stephan, et al.
Veröffentlicht: (2026) -
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023) -
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)