Multi-modal Relation Distillation for Unified 3D Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Huiqun, Bao, Yiping, Pan, Panwang, Li, Zeming, Liu, Xiao, Yang, Ruijie, Huang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
by: Mo, Wentao, et al.
Published: (2026)
by: Mo, Wentao, et al.
Published: (2026)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024)
by: Feng, Tuo, et al.
Published: (2024)
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
by: Liu, Peng
Published: (2021)
by: Liu, Peng
Published: (2021)
Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
by: Pan, Li, et al.
Published: (2024)
by: Pan, Li, et al.
Published: (2024)
Multi-modal Situated Reasoning in 3D Scenes
by: Linghu, Xiongkun, et al.
Published: (2024)
by: Linghu, Xiongkun, et al.
Published: (2024)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
by: Chen, Yiping, et al.
Published: (2026)
by: Chen, Yiping, et al.
Published: (2026)
UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
by: Zou, Jian, et al.
Published: (2023)
by: Zou, Jian, et al.
Published: (2023)
Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion
by: Li, Lijian
Published: (2025)
by: Li, Lijian
Published: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
by: Zhou, Zhaokun, et al.
Published: (2024)
by: Zhou, Zhaokun, et al.
Published: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
by: He, Yuting, et al.
Published: (2026)
by: He, Yuting, et al.
Published: (2026)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
by: Lei, Yinjie, et al.
Published: (2023)
by: Lei, Yinjie, et al.
Published: (2023)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection
by: Zhao, Haimei, et al.
Published: (2023)
by: Zhao, Haimei, et al.
Published: (2023)
Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke
by: Chen, Liren, et al.
Published: (2026)
by: Chen, Liren, et al.
Published: (2026)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR
by: Pathmanathan, Thenukan, et al.
Published: (2026)
by: Pathmanathan, Thenukan, et al.
Published: (2026)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023)
by: Lei, Youbo, et al.
Published: (2023)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
by: Cui, Jialei, et al.
Published: (2025)
by: Cui, Jialei, et al.
Published: (2025)
Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
by: Wang, Zhao-Yang, et al.
Published: (2025)
by: Wang, Zhao-Yang, et al.
Published: (2025)
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs
by: Lin, Yuhui, et al.
Published: (2026)
by: Lin, Yuhui, et al.
Published: (2026)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
by: Baby, Britty, et al.
Published: (2025)
by: Baby, Britty, et al.
Published: (2025)
Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and Distillation
by: Liu, Zhaoyu, et al.
Published: (2025)
by: Liu, Zhaoyu, et al.
Published: (2025)
TrajFlow: Multi-modal Motion Prediction via Flow Matching
by: Yan, Qi, et al.
Published: (2025)
by: Yan, Qi, et al.
Published: (2025)
Distilling Multi-view Diffusion Models into 3D Generators
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
by: Govindarajan, Hariprasath, et al.
Published: (2025)
by: Govindarajan, Hariprasath, et al.
Published: (2025)
HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation
by: Wang, Jiaxi, et al.
Published: (2023)
by: Wang, Jiaxi, et al.
Published: (2023)
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
by: Chen, Jiankang, et al.
Published: (2025)
by: Chen, Jiankang, et al.
Published: (2025)
Similar Items
-
Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation
by: Tan, Zhaorui, et al.
Published: (2024) -
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
by: Mo, Wentao, et al.
Published: (2026) -
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022) -
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025) -
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024)