TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yiyao, Zheng, Zhedong, Ziwei, Yu, Wang, Yaxiong, Tse, Tze Ho Elden, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025)
by: Han, Yi, et al.
Published: (2025)
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026)
by: Shatwell, David G., et al.
Published: (2026)
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
by: Xu, Kai, et al.
Published: (2024)
by: Xu, Kai, et al.
Published: (2024)
Improving Human Motion Plausibility with Body Momentum
by: Nguyen, Ha Linh, et al.
Published: (2025)
by: Nguyen, Ha Linh, et al.
Published: (2025)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2025)
by: Liu, Ruicong, et al.
Published: (2025)
GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement
by: Zheng, Linfang, et al.
Published: (2024)
by: Zheng, Linfang, et al.
Published: (2024)
A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point Clouds
by: Peng, Jizong, et al.
Published: (2025)
by: Peng, Jizong, et al.
Published: (2025)
Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery
by: Yang, Fengyuan, et al.
Published: (2024)
by: Yang, Fengyuan, et al.
Published: (2024)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025)
by: Tse, Tze Ho Elden, et al.
Published: (2025)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
by: Han, Guangyi, et al.
Published: (2025)
by: Han, Guangyi, et al.
Published: (2025)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
by: Liu, Lingyu, et al.
Published: (2025)
by: Liu, Lingyu, et al.
Published: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
by: Zhou, Yinan, et al.
Published: (2025)
by: Zhou, Yinan, et al.
Published: (2025)
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
SA-GS: Semantic-Aware Gaussian Splatting for Large Scene Reconstruction with Geometry Constrain
by: Xiong, Butian, et al.
Published: (2024)
by: Xiong, Butian, et al.
Published: (2024)
High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
by: Feng, Runyang, et al.
Published: (2025)
by: Feng, Runyang, et al.
Published: (2025)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
SCORP: Scene-Consistent Object Refinement via Proxy Generation and Tuning
by: Chen, Ziwei, et al.
Published: (2025)
by: Chen, Ziwei, et al.
Published: (2025)
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
by: Ji, Yihong, et al.
Published: (2025)
by: Ji, Yihong, et al.
Published: (2025)
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
by: Tian, Yongqi, et al.
Published: (2025)
by: Tian, Yongqi, et al.
Published: (2025)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
by: Ju, Hao, et al.
Published: (2025)
by: Ju, Hao, et al.
Published: (2025)
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
by: Cha, Junuk, et al.
Published: (2024)
by: Cha, Junuk, et al.
Published: (2024)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
by: Lian, Jingchun, et al.
Published: (2024)
by: Lian, Jingchun, et al.
Published: (2024)
Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling
by: Hao, Yuze, et al.
Published: (2024)
by: Hao, Yuze, et al.
Published: (2024)
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
by: Zhou, Bohan, et al.
Published: (2025)
by: Zhou, Bohan, et al.
Published: (2025)
Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection
by: Zhang, Ruiyang, et al.
Published: (2024)
by: Zhang, Ruiyang, et al.
Published: (2024)
Instilling Multi-round Thinking to Text-guided Image Generation
by: Zeng, Lidong, et al.
Published: (2024)
by: Zeng, Lidong, et al.
Published: (2024)
Harnessing Weak Pair Uncertainty for Text-based Person Search
by: Sun, Jintao, et al.
Published: (2026)
by: Sun, Jintao, et al.
Published: (2026)
HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data
by: Zhang, Mengqi, et al.
Published: (2024)
by: Zhang, Mengqi, et al.
Published: (2024)
Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
by: Zhang, Ruiyang, et al.
Published: (2024)
by: Zhang, Ruiyang, et al.
Published: (2024)
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
by: Zhang, Yuchen, et al.
Published: (2025)
by: Zhang, Yuchen, et al.
Published: (2025)
Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation
by: Xie, Xianghui, et al.
Published: (2023)
by: Xie, Xianghui, et al.
Published: (2023)
Similar Items
-
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024) -
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025) -
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026) -
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
by: Xu, Kai, et al.
Published: (2024) -
Improving Human Motion Plausibility with Body Momentum
by: Nguyen, Ha Linh, et al.
Published: (2025)