Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Mori, Erika, Qiu, Yue, Kataoka, Hirokatsu, Aoki, Yoshimitsu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025)
by: Sugiyama, Misora, et al.
Published: (2025)
Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
by: Torimi, Kohei, et al.
Published: (2025)
by: Torimi, Kohei, et al.
Published: (2025)
Industrial Synthetic Segment Pre-training
by: Mae, Shinichi, et al.
Published: (2025)
by: Mae, Shinichi, et al.
Published: (2025)
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025)
by: Kupyn, Orest, et al.
Published: (2025)
Data Collection-free Masked Video Modeling
by: Ishikawa, Yuchi, et al.
Published: (2024)
by: Ishikawa, Yuchi, et al.
Published: (2024)
Rethinking Image Super-Resolution from Training Data Perspectives
by: Ohtani, Go, et al.
Published: (2024)
by: Ohtani, Go, et al.
Published: (2024)
Human Action Recognition without Human
by: Kataoka, Hirokatsu, et al.
Published: (2016)
by: Kataoka, Hirokatsu, et al.
Published: (2016)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
by: Tateno, Masatoshi, et al.
Published: (2025)
by: Tateno, Masatoshi, et al.
Published: (2025)
BoundMatch: Boundary detection applied to semi-supervised segmentation
by: Ishikawa, Haruya, et al.
Published: (2025)
by: Ishikawa, Haruya, et al.
Published: (2025)
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
TAG: Guidance-free Open-Vocabulary Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion Models
by: Zhu, Peifei, et al.
Published: (2024)
by: Zhu, Peifei, et al.
Published: (2024)
Can masking background and object reduce static bias for zero-shot action recognition?
by: Fukuzawa, Takumi, et al.
Published: (2025)
by: Fukuzawa, Takumi, et al.
Published: (2025)
Pre-training with 3D Synthetic Data: Learning 3D Point Cloud Instance Segmentation from 3D Synthetic Scenes
by: Otsuka, Daichi, et al.
Published: (2025)
by: Otsuka, Daichi, et al.
Published: (2025)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
by: Xu, Zibo, et al.
Published: (2025)
by: Xu, Zibo, et al.
Published: (2025)
3D Human Scan With A Moving Event Camera
by: Kohyama, Kai, et al.
Published: (2024)
by: Kohyama, Kai, et al.
Published: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
by: Zou, Yuanhao, et al.
Published: (2025)
by: Zou, Yuanhao, et al.
Published: (2025)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
PowerCLIP: Powerset Alignment for Contrastive Pre-Training
by: Kawamura, Masaki, et al.
Published: (2025)
by: Kawamura, Masaki, et al.
Published: (2025)
Primitive Geometry Segment Pre-training for 3D Medical Image Segmentation
by: Tadokoro, Ryu, et al.
Published: (2024)
by: Tadokoro, Ryu, et al.
Published: (2024)
Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
by: Yoshihashi, Ryota, et al.
Published: (2023)
by: Yoshihashi, Ryota, et al.
Published: (2023)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
Top-down Activity Representation Learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
by: Xu, Quanxing, et al.
Published: (2026)
by: Xu, Quanxing, et al.
Published: (2026)
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
by: Tao, Qian, et al.
Published: (2025)
by: Tao, Qian, et al.
Published: (2025)
Iterative Event-based Motion Segmentation by Variational Contrast Maximization
by: Yamaki, Ryo, et al.
Published: (2025)
by: Yamaki, Ryo, et al.
Published: (2025)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering
by: Jin, Mengyuan, et al.
Published: (2026)
by: Jin, Mengyuan, et al.
Published: (2026)
Multi-object event graph representation learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Secrets of Event-Based Optical Flow
by: Shiba, Shintaro, et al.
Published: (2022)
by: Shiba, Shintaro, et al.
Published: (2022)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
by: Meng, Yiran, et al.
Published: (2025)
by: Meng, Yiran, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
by: Zemskova, Tatiana, et al.
Published: (2026)
by: Zemskova, Tatiana, et al.
Published: (2026)
FDIF: Formula-Driven supervised Learning with Implicit Functions for 3D Medical Image Segmentation
by: Yamamoto, Yukinori, et al.
Published: (2026)
by: Yamamoto, Yukinori, et al.
Published: (2026)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
Similar Items
-
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025) -
Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
by: Torimi, Kohei, et al.
Published: (2025) -
Industrial Synthetic Segment Pre-training
by: Mae, Shinichi, et al.
Published: (2025) -
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025) -
Data Collection-free Masked Video Modeling
by: Ishikawa, Yuchi, et al.
Published: (2024)