GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering
Fuente:
arXiv
Guardado en:
| Autores principales: | Brilli, Dionysia Danai, Mallis, Dimitrios, Pitsikalis, Vassilis, Maragos, Petros |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AIris: An AI-powered Wearable Assistive Device for the Visually Impaired
por: Brilli, Dionysia Danai, et al.
Publicado: (2024)
por: Brilli, Dionysia Danai, et al.
Publicado: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
por: Jia, Ziheng, et al.
Publicado: (2024)
por: Jia, Ziheng, et al.
Publicado: (2024)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
por: Tran, Duong T., et al.
Publicado: (2025)
por: Tran, Duong T., et al.
Publicado: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
por: Min, Juhong, et al.
Publicado: (2024)
por: Min, Juhong, et al.
Publicado: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
por: Meng, Yiran, et al.
Publicado: (2025)
por: Meng, Yiran, et al.
Publicado: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
por: Drago, Mauro Orazio, et al.
Publicado: (2025)
por: Drago, Mauro Orazio, et al.
Publicado: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
por: Vasileiou, Vasiliki, et al.
Publicado: (2026)
por: Vasileiou, Vasiliki, et al.
Publicado: (2026)
Multi-task Learning For Joint Action and Gesture Recognition
por: Spathis, Konstantinos, et al.
Publicado: (2025)
por: Spathis, Konstantinos, et al.
Publicado: (2025)
TropNNC: Structured Neural Network Compression Using Tropical Geometry
por: Fotopoulos, Konstantinos, et al.
Publicado: (2024)
por: Fotopoulos, Konstantinos, et al.
Publicado: (2024)
Pre-training for Action Recognition with Automatically Generated Fractal Datasets
por: Svyezhentsev, Davyd, et al.
Publicado: (2024)
por: Svyezhentsev, Davyd, et al.
Publicado: (2024)
Mushroom Segmentation and 3D Pose Estimation from Point Clouds using Fully Convolutional Geometric Features and Implicit Pose Encoding
por: Retsinas, George, et al.
Publicado: (2024)
por: Retsinas, George, et al.
Publicado: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2025)
por: Zhou, Sheng, et al.
Publicado: (2025)
BERT-VQA: Visual Question Answering on Plots
por: Vu, Tai, et al.
Publicado: (2025)
por: Vu, Tai, et al.
Publicado: (2025)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
por: Zhang, Xiaoman, et al.
Publicado: (2023)
por: Zhang, Xiaoman, et al.
Publicado: (2023)
CommVQA: Situating Visual Question Answering in Communicative Contexts
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
por: Nguyen, Hai-Dang, et al.
Publicado: (2025)
por: Nguyen, Hai-Dang, et al.
Publicado: (2025)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
por: Al-Mohannadi, Aisha, et al.
Publicado: (2026)
por: Al-Mohannadi, Aisha, et al.
Publicado: (2026)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
por: Chaybouti, Sofian, et al.
Publicado: (2025)
por: Chaybouti, Sofian, et al.
Publicado: (2025)
Only-Style: Stylistic Consistency in Image Generation without Content Leakage
por: Aravanis, Tilemachos, et al.
Publicado: (2025)
por: Aravanis, Tilemachos, et al.
Publicado: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
por: Fernando, Basura, et al.
Publicado: (2025)
por: Fernando, Basura, et al.
Publicado: (2025)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
por: Madaka, Madhuri Latha, et al.
Publicado: (2025)
por: Madaka, Madhuri Latha, et al.
Publicado: (2025)
CogStream: Context-guided Streaming Video Question Answering
por: Zhao, Zicheng, et al.
Publicado: (2025)
por: Zhao, Zicheng, et al.
Publicado: (2025)
Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data
por: Glytsos, Marios, et al.
Publicado: (2025)
por: Glytsos, Marios, et al.
Publicado: (2025)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
por: Fan, Sunqi, et al.
Publicado: (2025)
por: Fan, Sunqi, et al.
Publicado: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
por: Vu, Sinh Trong, et al.
Publicado: (2025)
por: Vu, Sinh Trong, et al.
Publicado: (2025)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
por: Liao, Zhaohe, et al.
Publicado: (2024)
por: Liao, Zhaohe, et al.
Publicado: (2024)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
por: He, Runlong, et al.
Publicado: (2024)
por: He, Runlong, et al.
Publicado: (2024)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
por: Li, Zhifei, et al.
Publicado: (2026)
por: Li, Zhifei, et al.
Publicado: (2026)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
por: Wang, Jialou, et al.
Publicado: (2024)
por: Wang, Jialou, et al.
Publicado: (2024)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
por: Zhang, Chengyi, et al.
Publicado: (2026)
por: Zhang, Chengyi, et al.
Publicado: (2026)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
por: He, Zhixian, et al.
Publicado: (2024)
por: He, Zhixian, et al.
Publicado: (2024)
TransCAD: A Hierarchical Transformer for CAD Sequence Inference from Point Clouds
por: Dupont, Elona, et al.
Publicado: (2024)
por: Dupont, Elona, et al.
Publicado: (2024)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
por: Zemskova, Tatiana, et al.
Publicado: (2026)
por: Zemskova, Tatiana, et al.
Publicado: (2026)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
por: Guan, Runwei, et al.
Publicado: (2025)
por: Guan, Runwei, et al.
Publicado: (2025)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
por: Hu, Rongsheng, et al.
Publicado: (2026)
por: Hu, Rongsheng, et al.
Publicado: (2026)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
por: Chen, Pingyi, et al.
Publicado: (2024)
por: Chen, Pingyi, et al.
Publicado: (2024)
Optimal Transport for Handwritten Text Recognition in a Low-Resource Regime
por: Wraight, Petros Georgoulas, et al.
Publicado: (2025)
por: Wraight, Petros Georgoulas, et al.
Publicado: (2025)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
por: Chatzichristodoulou, Georgios, et al.
Publicado: (2025)
por: Chatzichristodoulou, Georgios, et al.
Publicado: (2025)
Ejemplares similares
-
AIris: An AI-powered Wearable Assistive Device for the Visually Impaired
por: Brilli, Dionysia Danai, et al.
Publicado: (2024) -
VQA$^2$: Visual Question Answering for Video Quality Assessment
por: Jia, Ziheng, et al.
Publicado: (2024) -
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
por: Tran, Duong T., et al.
Publicado: (2025) -
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
por: Min, Juhong, et al.
Publicado: (2024) -
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
por: Meng, Yiran, et al.
Publicado: (2025)