Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Ruoyue, Inoue, Nakamasa, Shinoda, Koichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
von: Ukai, Mahiro, et al.
Veröffentlicht: (2025)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2025)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
VoQA: Visual-only Question Answering
von: An, Jianing, et al.
Veröffentlicht: (2025)
von: An, Jianing, et al.
Veröffentlicht: (2025)
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
von: Ahir, Param, et al.
Veröffentlicht: (2023)
von: Ahir, Param, et al.
Veröffentlicht: (2023)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
von: Chen, Pingyi, et al.
Veröffentlicht: (2024)
von: Chen, Pingyi, et al.
Veröffentlicht: (2024)
D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
von: Rahimi, Amir, et al.
Veröffentlicht: (2023)
von: Rahimi, Amir, et al.
Veröffentlicht: (2023)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
Multi-Point Positional Insertion Tuning for Small Object Detection
von: Goto, Kanoko, et al.
Veröffentlicht: (2024)
von: Goto, Kanoko, et al.
Veröffentlicht: (2024)
Free Form Medical Visual Question Answering in Radiology
von: Narayanan, Abhishek, et al.
Veröffentlicht: (2024)
von: Narayanan, Abhishek, et al.
Veröffentlicht: (2024)
Referring Expression Comprehension for Small Objects
von: Goto, Kanoko, et al.
Veröffentlicht: (2025)
von: Goto, Kanoko, et al.
Veröffentlicht: (2025)
Saliency Guided Longitudinal Medical Visual Question Answering
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
von: Özdemir, Övgü, et al.
Veröffentlicht: (2024)
von: Özdemir, Övgü, et al.
Veröffentlicht: (2024)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design
von: Tang, Wenxin, et al.
Veröffentlicht: (2025)
von: Tang, Wenxin, et al.
Veröffentlicht: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
von: Marouf, Imad Eddine, et al.
Veröffentlicht: (2025)
von: Marouf, Imad Eddine, et al.
Veröffentlicht: (2025)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
Location-Aware Pretraining for Medical Difference Visual Question Answering
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
IIU: Independent Inference Units for Knowledge-based Visual Question Answering
von: Li, Yili, et al.
Veröffentlicht: (2024)
von: Li, Yili, et al.
Veröffentlicht: (2024)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
von: Jain, Riddhi, et al.
Veröffentlicht: (2025)
von: Jain, Riddhi, et al.
Veröffentlicht: (2025)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
LingoQA: Visual Question Answering for Autonomous Driving
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
von: Yokomizo, Hisayuki, et al.
Veröffentlicht: (2026)
von: Yokomizo, Hisayuki, et al.
Veröffentlicht: (2026)
CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation
von: Doris, Anna C., et al.
Veröffentlicht: (2025)
von: Doris, Anna C., et al.
Veröffentlicht: (2025)
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024) -
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025) -
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025) -
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
von: Zhang, Junkai, et al.
Veröffentlicht: (2025) -
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
von: Ukai, Mahiro, et al.
Veröffentlicht: (2025)