Training-Free Dense Hand Contact Estimation with Multi-Modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Daniel Sungho, Lee, Kyoung Mu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Dense Hand Contact Estimation from Imbalanced Data
by: Jung, Daniel Sungho, et al.
Published: (2025)
by: Jung, Daniel Sungho, et al.
Published: (2025)
Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
by: Jung, Daniel Sungho, et al.
Published: (2025)
by: Jung, Daniel Sungho, et al.
Published: (2025)
Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer
by: Nam, Hyeongjin, et al.
Published: (2024)
by: Nam, Hyeongjin, et al.
Published: (2024)
Learning Human-Object Interaction for 3D Human Pose Estimation from LiDAR Point Clouds
by: Jung, Daniel Sungho, et al.
Published: (2026)
by: Jung, Daniel Sungho, et al.
Published: (2026)
SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models
by: Lee, Jaerin, et al.
Published: (2024)
by: Lee, Jaerin, et al.
Published: (2024)
TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
by: Nam, Hyeongjin, et al.
Published: (2026)
by: Nam, Hyeongjin, et al.
Published: (2026)
StarryGazer: Leveraging Monocular Depth Estimation Models for Domain-Agnostic Single Depth Image Completion
by: Hong, Sangmin, et al.
Published: (2025)
by: Hong, Sangmin, et al.
Published: (2025)
Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images
by: Lee, Kanggeon, et al.
Published: (2026)
by: Lee, Kanggeon, et al.
Published: (2026)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
Particle Diffusion Matching: Random Walk Correspondence Search for the Alignment of Standard and Ultra-Widefield Fundus Images
by: Lee, Kanggeon, et al.
Published: (2026)
by: Lee, Kanggeon, et al.
Published: (2026)
Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization
by: Lee, Jaerin, et al.
Published: (2026)
by: Lee, Jaerin, et al.
Published: (2026)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
by: Cheng, Wencan, et al.
Published: (2026)
by: Cheng, Wencan, et al.
Published: (2026)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
by: Lin, Nie, et al.
Published: (2024)
by: Lin, Nie, et al.
Published: (2024)
GS-Blur: A 3D Scene-Based Dataset for Realistic Image Deblurring
by: Lee, Dongwoo, et al.
Published: (2024)
by: Lee, Dongwoo, et al.
Published: (2024)
GLOS: Sign Language Generation with Temporally Aligned Gloss-Level Conditioning
by: Lee, Taeryung, et al.
Published: (2025)
by: Lee, Taeryung, et al.
Published: (2025)
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
by: Seifi, Soroush, et al.
Published: (2025)
by: Seifi, Soroush, et al.
Published: (2025)
MatRes: Zero-Shot Test-Time Model Adaptation for Simultaneous Matching and Restoration
by: Lee, Kanggeon, et al.
Published: (2026)
by: Lee, Kanggeon, et al.
Published: (2026)
NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object Interaction
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
Exploiting Diffusion Prior for Task-driven Image Restoration
by: Kim, Jaeha, et al.
Published: (2025)
by: Kim, Jaeha, et al.
Published: (2025)
Beyond Image Super-Resolution for Image Recognition with Task-Driven Perceptual Loss
by: Kim, Jaeha, et al.
Published: (2024)
by: Kim, Jaeha, et al.
Published: (2024)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
by: Imran, Muhammad, et al.
Published: (2025)
by: Imran, Muhammad, et al.
Published: (2025)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
by: Sarowar, Md Selim, et al.
Published: (2025)
by: Sarowar, Md Selim, et al.
Published: (2025)
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
by: Liu, Jie, et al.
Published: (2024)
by: Liu, Jie, et al.
Published: (2024)
TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts
by: Ktistakis, Sophokles, et al.
Published: (2026)
by: Ktistakis, Sophokles, et al.
Published: (2026)
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
by: Oh, Yeongtak, et al.
Published: (2025)
by: Oh, Yeongtak, et al.
Published: (2025)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
ODGS: 3D Scene Reconstruction from Omnidirectional Images with 3D Gaussian Splattings
by: Lee, Suyoung, et al.
Published: (2024)
by: Lee, Suyoung, et al.
Published: (2024)
MEIL-NeRF: Memory-Efficient Incremental Learning of Neural Radiance Fields
by: Chung, Jaeyoung, et al.
Published: (2022)
by: Chung, Jaeyoung, et al.
Published: (2022)
DeblurGS: Gaussian Splatting for Camera Motion Blur
by: Oh, Jeongtaek, et al.
Published: (2024)
by: Oh, Jeongtaek, et al.
Published: (2024)
OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned Synthesis
by: Cha, Junuk, et al.
Published: (2026)
by: Cha, Junuk, et al.
Published: (2026)
Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics
by: Cho, Woojin, et al.
Published: (2024)
by: Cho, Woojin, et al.
Published: (2024)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
by: Zhu, Yabin, et al.
Published: (2026)
by: Zhu, Yabin, et al.
Published: (2026)
Auto-regressive transformation for image alignment
by: Lee, Kanggeon, et al.
Published: (2025)
by: Lee, Kanggeon, et al.
Published: (2025)
Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
by: Bonat, Laurence, et al.
Published: (2026)
by: Bonat, Laurence, et al.
Published: (2026)
Similar Items
-
Learning Dense Hand Contact Estimation from Imbalanced Data
by: Jung, Daniel Sungho, et al.
Published: (2025) -
Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
by: Jung, Daniel Sungho, et al.
Published: (2025) -
Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer
by: Nam, Hyeongjin, et al.
Published: (2024) -
Learning Human-Object Interaction for 3D Human Pose Estimation from LiDAR Point Clouds
by: Jung, Daniel Sungho, et al.
Published: (2026) -
SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models
by: Lee, Jaerin, et al.
Published: (2024)