TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhihao, Cao, Shengcao, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
by: Lu, Yuxiang, et al.
Published: (2024)
by: Lu, Yuxiang, et al.
Published: (2024)
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
by: Guo, Hanzhong, et al.
Published: (2024)
by: Guo, Hanzhong, et al.
Published: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024)
by: Wang, Qijie, et al.
Published: (2024)
IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
by: Zhong, Xiaojing, et al.
Published: (2025)
by: Zhong, Xiaojing, et al.
Published: (2025)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
by: Xing, Jiazheng, et al.
Published: (2024)
by: Xing, Jiazheng, et al.
Published: (2024)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
by: Zhu, Yabin, et al.
Published: (2026)
by: Zhu, Yabin, et al.
Published: (2026)
SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking
by: Hou, Xiaojun, et al.
Published: (2024)
by: Hou, Xiaojun, et al.
Published: (2024)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
by: Zhao, Yumiao, et al.
Published: (2024)
by: Zhao, Yumiao, et al.
Published: (2024)
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
by: Li, Fan, et al.
Published: (2025)
by: Li, Fan, et al.
Published: (2025)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Direction-Aware Hybrid Representation Learning for 3D Hand Pose and Shape Estimation
by: Liu, Shiyong, et al.
Published: (2025)
by: Liu, Shiyong, et al.
Published: (2025)
Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter
by: Zha, Yaohua, et al.
Published: (2025)
by: Zha, Yaohua, et al.
Published: (2025)
OmniTry: Virtual Try-On Anything without Masks
by: Feng, Yutong, et al.
Published: (2025)
by: Feng, Yutong, et al.
Published: (2025)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
VSFormer: Mining Correlations in Flexible View Set for Multi-view 3D Shape Understanding
by: Sun, Hongyu, et al.
Published: (2024)
by: Sun, Hongyu, et al.
Published: (2024)
Shape-Guided Clothing Warping for Virtual Try-On
by: Han, Xiaoyu, et al.
Published: (2025)
by: Han, Xiaoyu, et al.
Published: (2025)
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
by: Xu, Boyue, et al.
Published: (2025)
by: Xu, Boyue, et al.
Published: (2025)
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
by: Wang, Aowen, et al.
Published: (2025)
by: Wang, Aowen, et al.
Published: (2025)
TriFusion-SR: Joint Tri-Modal Medical Image Fusion and SR
by: Dharejo, Fayaz Ali, et al.
Published: (2026)
by: Dharejo, Fayaz Ali, et al.
Published: (2026)
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
TC-GS: Tri-plane based compression for 3D Gaussian Splatting
by: Wang, Taorui, et al.
Published: (2025)
by: Wang, Taorui, et al.
Published: (2025)
Symmetry Understanding of 3D Shapes via Chirality Disentanglement
by: Wang, Weikang, et al.
Published: (2025)
by: Wang, Weikang, et al.
Published: (2025)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
by: Wang, Yiqun, et al.
Published: (2024)
by: Wang, Yiqun, et al.
Published: (2024)
MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation
by: Chen, Zhuangzhuang, et al.
Published: (2025)
by: Chen, Zhuangzhuang, et al.
Published: (2025)
Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos
by: Zhang, Gangjian, et al.
Published: (2026)
by: Zhang, Gangjian, et al.
Published: (2026)
COM3D: Leveraging Cross-View Correspondence and Cross-Modal Mining for 3D Retrieval
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
by: Xiong, Zeren, et al.
Published: (2025)
by: Xiong, Zeren, et al.
Published: (2025)
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
by: Reza, Sakib, et al.
Published: (2025)
by: Reza, Sakib, et al.
Published: (2025)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
Flexible 3D Lane Detection by Hierarchical Shape MatchingFlexible 3D Lane Detection by Hierarchical Shape Matching
by: Guan, Zhihao, et al.
Published: (2024)
by: Guan, Zhihao, et al.
Published: (2024)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
by: Qi, Zekun, et al.
Published: (2024)
by: Qi, Zekun, et al.
Published: (2024)
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
Similar Items
-
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
by: Lu, Yuxiang, et al.
Published: (2024) -
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
by: Guo, Hanzhong, et al.
Published: (2024) -
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024) -
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024) -
IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
by: Zhong, Xiaojing, et al.
Published: (2025)