HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Sayem, MD Khalequzzaman Chowdhury, Chowdhury, Mubarrat Tajoar, Tiruneh, Yihalem Yimolal, Khan, Muneeb A., Ali, Muhammad Salman, Bhattarai, Binod, Baek, Seungryul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating Objects
di: Ismayilzada, Elkhan, et al.
Pubblicazione: (2025)
di: Ismayilzada, Elkhan, et al.
Pubblicazione: (2025)
THOM: Generating Physically Plausible Hand-Object Meshes From Text
di: Jeong, Uyoung, et al.
Pubblicazione: (2026)
di: Jeong, Uyoung, et al.
Pubblicazione: (2026)
SDDGR: Stable Diffusion-based Deep Generative Replay for Class Incremental Object Detection
di: Kim, Junsu, et al.
Pubblicazione: (2024)
di: Kim, Junsu, et al.
Pubblicazione: (2024)
GeoBotsVR: A Robotics Learning Game for Beginners with Hands-on Learning Simulation
di: Mubarrat, Syed T.
Pubblicazione: (2024)
di: Mubarrat, Syed T.
Pubblicazione: (2024)
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
di: Cha, Junuk, et al.
Pubblicazione: (2024)
di: Cha, Junuk, et al.
Pubblicazione: (2024)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
di: Tateno, Masatoshi, et al.
Pubblicazione: (2025)
di: Tateno, Masatoshi, et al.
Pubblicazione: (2025)
Local K-Similarity Constraint for Federated Learning with Label Noise
di: Amgain, Sanskar, et al.
Pubblicazione: (2025)
di: Amgain, Sanskar, et al.
Pubblicazione: (2025)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
Input Invex Neural Network
di: Sapkota, Suman, et al.
Pubblicazione: (2021)
di: Sapkota, Suman, et al.
Pubblicazione: (2021)
Dimension Mixer: Group Mixing of Input Dimensions for Efficient Function Approximation
di: Sapkota, Suman, et al.
Pubblicazione: (2023)
di: Sapkota, Suman, et al.
Pubblicazione: (2023)
ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning
di: Chowdhury, MD Thamed Bin Zaman, et al.
Pubblicazione: (2025)
di: Chowdhury, MD Thamed Bin Zaman, et al.
Pubblicazione: (2025)
ARISTO Hand: Sensing-Driven Distal Hyperextension for Fine-Grained Manipulation
di: Kim, Aaron, et al.
Pubblicazione: (2026)
di: Kim, Aaron, et al.
Pubblicazione: (2026)
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
di: Gan, Qijun, et al.
Pubblicazione: (2024)
di: Gan, Qijun, et al.
Pubblicazione: (2024)
Beyond Words: ESC‐Net Revolutionizes VQA by Elevating Visual Features and Defying Language Priors
di: Souvik Chowdhury, et al.
Pubblicazione: (2024)
di: Souvik Chowdhury, et al.
Pubblicazione: (2024)
#InSafeHands: not about the money
Pubblicazione: (2026)
Pubblicazione: (2026)
Can Synthetic Images Conquer Forgetting? Beyond Unexplored Doubts in Few-Shot Class-Incremental Learning
di: Kim, Junsu, et al.
Pubblicazione: (2025)
di: Kim, Junsu, et al.
Pubblicazione: (2025)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
di: Pei, Baoqi, et al.
Pubblicazione: (2025)
di: Pei, Baoqi, et al.
Pubblicazione: (2025)
Abduction of Domain Relationships from Data for VQA
di: Chowdhury, Al Mehdi Saadat, et al.
Pubblicazione: (2025)
di: Chowdhury, Al Mehdi Saadat, et al.
Pubblicazione: (2025)
On the Monotonicity of Information Aging
di: Shisher, MD Kamran Chowdhury, et al.
Pubblicazione: (2024)
di: Shisher, MD Kamran Chowdhury, et al.
Pubblicazione: (2024)
The Hand Behind the Invisible Hand
di: Mittermaier, Karl
Pubblicazione: (2020)
di: Mittermaier, Karl
Pubblicazione: (2020)
Hand in Hand: Technology Inclusion.
Pubblicazione: (1994)
Pubblicazione: (1994)
Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation
di: Duangprom, Krit, et al.
Pubblicazione: (2025)
di: Duangprom, Krit, et al.
Pubblicazione: (2025)
An Empirical Analysis of Fine-Tuning Large Language Models on Bioinformatics Literature: PRSGPT and BioStarsGPT
di: Muneeb, Muhammad, et al.
Pubblicazione: (2025)
di: Muneeb, Muhammad, et al.
Pubblicazione: (2025)
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
di: Fan, Zicong, et al.
Pubblicazione: (2024)
di: Fan, Zicong, et al.
Pubblicazione: (2024)
MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
di: Saleem, Muhammad Usama, et al.
Pubblicazione: (2024)
di: Saleem, Muhammad Usama, et al.
Pubblicazione: (2024)
On the Utility of 3D Hand Poses for Action Recognition
di: Shamil, Md Salman, et al.
Pubblicazione: (2024)
di: Shamil, Md Salman, et al.
Pubblicazione: (2024)
Beyond Synthetic Replays: Turning Diffusion Features into Few-Shot Class-Incremental Learning Knowledge
di: Kim, Junsu, et al.
Pubblicazione: (2025)
di: Kim, Junsu, et al.
Pubblicazione: (2025)
Design and Application of Multimodal Large Language Model Based System for End to End Automation of Accident Dataset Generation
di: Chowdhury, MD Thamed Bin Zaman, et al.
Pubblicazione: (2025)
di: Chowdhury, MD Thamed Bin Zaman, et al.
Pubblicazione: (2025)
Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
di: Zheng, Hao, et al.
Pubblicazione: (2026)
di: Zheng, Hao, et al.
Pubblicazione: (2026)
Attention-based Estimation and Prediction of Human Intent to augment Haptic Glove aided Control of Robotic Hand
di: Ahmed, Muneeb, et al.
Pubblicazione: (2021)
di: Ahmed, Muneeb, et al.
Pubblicazione: (2021)
Attention-Guided Dual-Stream Learning for Group Engagement Recognition: Fusing Transformer-Encoded Motion Dynamics with Scene Context via Adaptive Gating
di: Chowdhury, Saniah Kayenat, et al.
Pubblicazione: (2026)
di: Chowdhury, Saniah Kayenat, et al.
Pubblicazione: (2026)
Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling
di: Hao, Yuze, et al.
Pubblicazione: (2024)
di: Hao, Yuze, et al.
Pubblicazione: (2024)
Hand in Hand: Media Literacy and Internet Safety
di: Gallagher, Frank
Pubblicazione: (2011)
di: Gallagher, Frank
Pubblicazione: (2011)
Writing and Word Processing Go Hand in Hand.
di: Casella, Vicki
Pubblicazione: (1989)
di: Casella, Vicki
Pubblicazione: (1989)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
di: Bodur, Rumeysa, et al.
Pubblicazione: (2024)
di: Bodur, Rumeysa, et al.
Pubblicazione: (2024)
VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embedding
di: Liang, Yujie, et al.
Pubblicazione: (2024)
di: Liang, Yujie, et al.
Pubblicazione: (2024)
Two Hands Are Better Than One: Resolving Hand to Hand Intersections via Occupancy Networks
di: Ivashechkin, Maksym, et al.
Pubblicazione: (2024)
di: Ivashechkin, Maksym, et al.
Pubblicazione: (2024)
Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
di: Kim, Taehoon, et al.
Pubblicazione: (2025)
di: Kim, Taehoon, et al.
Pubblicazione: (2025)
GenDexHand: Generative Simulation for Dexterous Hands
di: Chen, Feng, et al.
Pubblicazione: (2025)
di: Chen, Feng, et al.
Pubblicazione: (2025)
Leveraging deep neural networks to uncover unprecedented levels of precision in the diagnosis of hair and scalp disorders
di: Mohammad Sayem Chowdhury, et al.
Pubblicazione: (2024)
di: Mohammad Sayem Chowdhury, et al.
Pubblicazione: (2024)
Documenti analoghi
-
QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating Objects
di: Ismayilzada, Elkhan, et al.
Pubblicazione: (2025) -
THOM: Generating Physically Plausible Hand-Object Meshes From Text
di: Jeong, Uyoung, et al.
Pubblicazione: (2026) -
SDDGR: Stable Diffusion-based Deep Generative Replay for Class Incremental Object Detection
di: Kim, Junsu, et al.
Pubblicazione: (2024) -
GeoBotsVR: A Robotics Learning Game for Beginners with Hands-on Learning Simulation
di: Mubarrat, Syed T.
Pubblicazione: (2024) -
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
di: Cha, Junuk, et al.
Pubblicazione: (2024)