TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhihao, Cao, Shengcao, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
von: Guo, Hanzhong, et al.
Veröffentlicht: (2024)
von: Guo, Hanzhong, et al.
Veröffentlicht: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
von: Zhong, Xiaojing, et al.
Veröffentlicht: (2025)
von: Zhong, Xiaojing, et al.
Veröffentlicht: (2025)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
von: Zhang, Jusheng, et al.
Veröffentlicht: (2026)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2026)
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
von: Xing, Jiazheng, et al.
Veröffentlicht: (2024)
von: Xing, Jiazheng, et al.
Veröffentlicht: (2024)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking
von: Hou, Xiaojun, et al.
Veröffentlicht: (2024)
von: Hou, Xiaojun, et al.
Veröffentlicht: (2024)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
von: Li, Fan, et al.
Veröffentlicht: (2025)
von: Li, Fan, et al.
Veröffentlicht: (2025)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Direction-Aware Hybrid Representation Learning for 3D Hand Pose and Shape Estimation
von: Liu, Shiyong, et al.
Veröffentlicht: (2025)
von: Liu, Shiyong, et al.
Veröffentlicht: (2025)
Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter
von: Zha, Yaohua, et al.
Veröffentlicht: (2025)
von: Zha, Yaohua, et al.
Veröffentlicht: (2025)
OmniTry: Virtual Try-On Anything without Masks
von: Feng, Yutong, et al.
Veröffentlicht: (2025)
von: Feng, Yutong, et al.
Veröffentlicht: (2025)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
VSFormer: Mining Correlations in Flexible View Set for Multi-view 3D Shape Understanding
von: Sun, Hongyu, et al.
Veröffentlicht: (2024)
von: Sun, Hongyu, et al.
Veröffentlicht: (2024)
Shape-Guided Clothing Warping for Virtual Try-On
von: Han, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Han, Xiaoyu, et al.
Veröffentlicht: (2025)
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
von: Wang, Aowen, et al.
Veröffentlicht: (2025)
von: Wang, Aowen, et al.
Veröffentlicht: (2025)
TriFusion-SR: Joint Tri-Modal Medical Image Fusion and SR
von: Dharejo, Fayaz Ali, et al.
Veröffentlicht: (2026)
von: Dharejo, Fayaz Ali, et al.
Veröffentlicht: (2026)
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
TC-GS: Tri-plane based compression for 3D Gaussian Splatting
von: Wang, Taorui, et al.
Veröffentlicht: (2025)
von: Wang, Taorui, et al.
Veröffentlicht: (2025)
Symmetry Understanding of 3D Shapes via Chirality Disentanglement
von: Wang, Weikang, et al.
Veröffentlicht: (2025)
von: Wang, Weikang, et al.
Veröffentlicht: (2025)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
von: Wang, Yiqun, et al.
Veröffentlicht: (2024)
von: Wang, Yiqun, et al.
Veröffentlicht: (2024)
MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation
von: Chen, Zhuangzhuang, et al.
Veröffentlicht: (2025)
von: Chen, Zhuangzhuang, et al.
Veröffentlicht: (2025)
Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos
von: Zhang, Gangjian, et al.
Veröffentlicht: (2026)
von: Zhang, Gangjian, et al.
Veröffentlicht: (2026)
COM3D: Leveraging Cross-View Correspondence and Cross-Modal Mining for 3D Retrieval
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
von: Xiong, Zeren, et al.
Veröffentlicht: (2025)
von: Xiong, Zeren, et al.
Veröffentlicht: (2025)
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
von: Reza, Sakib, et al.
Veröffentlicht: (2025)
von: Reza, Sakib, et al.
Veröffentlicht: (2025)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
von: Wu, Zehuan, et al.
Veröffentlicht: (2024)
von: Wu, Zehuan, et al.
Veröffentlicht: (2024)
Flexible 3D Lane Detection by Hierarchical Shape MatchingFlexible 3D Lane Detection by Hierarchical Shape Matching
von: Guan, Zhihao, et al.
Veröffentlicht: (2024)
von: Guan, Zhihao, et al.
Veröffentlicht: (2024)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
von: Qi, Zekun, et al.
Veröffentlicht: (2024)
von: Qi, Zekun, et al.
Veröffentlicht: (2024)
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024) -
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
von: Guo, Hanzhong, et al.
Veröffentlicht: (2024) -
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
von: Cao, Shengcao, et al.
Veröffentlicht: (2024) -
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
von: Wang, Qijie, et al.
Veröffentlicht: (2024) -
IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
von: Zhong, Xiaojing, et al.
Veröffentlicht: (2025)