HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Tateno, Masatoshi, Kato, Gido, Kataoka, Hirokatsu, Sato, Yoichi, Yagi, Takuma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Multiple Object States from Actions via Large Language Models
by: Tateno, Masatoshi, et al.
Published: (2024)
by: Tateno, Masatoshi, et al.
Published: (2024)
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
by: Yagi, Takuma, et al.
Published: (2024)
by: Yagi, Takuma, et al.
Published: (2024)
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025)
by: Sugiyama, Misora, et al.
Published: (2025)
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025)
by: Kupyn, Orest, et al.
Published: (2025)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering
by: Mori, Erika, et al.
Published: (2025)
by: Mori, Erika, et al.
Published: (2025)
DyTact: Capturing Dynamic Contacts in Hand-Object Manipulation
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
by: Fan, Zicong, et al.
Published: (2024)
by: Fan, Zicong, et al.
Published: (2024)
Industrial Synthetic Segment Pre-training
by: Mae, Shinichi, et al.
Published: (2025)
by: Mae, Shinichi, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
Human Action Recognition without Human
by: Kataoka, Hirokatsu, et al.
Published: (2016)
by: Kataoka, Hirokatsu, et al.
Published: (2016)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion Models
by: Zhu, Peifei, et al.
Published: (2024)
by: Zhu, Peifei, et al.
Published: (2024)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
by: Hu, BoCheng, et al.
Published: (2026)
by: Hu, BoCheng, et al.
Published: (2026)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Can masking background and object reduce static bias for zero-shot action recognition?
by: Fukuzawa, Takumi, et al.
Published: (2025)
by: Fukuzawa, Takumi, et al.
Published: (2025)
Pre-training with 3D Synthetic Data: Learning 3D Point Cloud Instance Segmentation from 3D Synthetic Scenes
by: Otsuka, Daichi, et al.
Published: (2025)
by: Otsuka, Daichi, et al.
Published: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
by: Lee, Jihyun, et al.
Published: (2026)
by: Lee, Jihyun, et al.
Published: (2026)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
by: Furuta, Ryosuke, et al.
Published: (2023)
by: Furuta, Ryosuke, et al.
Published: (2023)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
by: Banno, Tatsuro, et al.
Published: (2025)
by: Banno, Tatsuro, et al.
Published: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2024)
by: Liu, Ruicong, et al.
Published: (2024)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
by: Yu, Suhao, et al.
Published: (2025)
by: Yu, Suhao, et al.
Published: (2025)
Primitive Geometry Segment Pre-training for 3D Medical Image Segmentation
by: Tadokoro, Ryu, et al.
Published: (2024)
by: Tadokoro, Ryu, et al.
Published: (2024)
Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
by: Yoshihashi, Ryota, et al.
Published: (2023)
by: Yoshihashi, Ryota, et al.
Published: (2023)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
PowerCLIP: Powerset Alignment for Contrastive Pre-Training
by: Kawamura, Masaki, et al.
Published: (2025)
by: Kawamura, Masaki, et al.
Published: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
by: Lin, Nie, et al.
Published: (2024)
by: Lin, Nie, et al.
Published: (2024)
MoireDB: Formula-generated Interference-fringe Image Dataset
by: Matsuo, Yuto, et al.
Published: (2025)
by: Matsuo, Yuto, et al.
Published: (2025)
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
by: Ouyang, Liangyang, et al.
Published: (2026)
by: Ouyang, Liangyang, et al.
Published: (2026)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Similar Items
-
Learning Multiple Object States from Actions via Large Language Models
by: Tateno, Masatoshi, et al.
Published: (2024) -
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
by: Yagi, Takuma, et al.
Published: (2024) -
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025) -
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025) -
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)