MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiaze, Wang, Yi, Guo, Ziyu, Zhang, Renrui, Zhou, Donghao, Chen, Guangyong, Liu, Anfeng, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
by: Zhou, Donghao, et al.
Published: (2024)
by: Zhou, Donghao, et al.
Published: (2024)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
by: Guo, Ziyu, et al.
Published: (2024)
by: Guo, Ziyu, et al.
Published: (2024)
SFANet: Spatial-Frequency Attention Network for Weather Forecasting
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
by: Zhou, Donghao, et al.
Published: (2023)
by: Zhou, Donghao, et al.
Published: (2023)
DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
by: Gao, Jialin, et al.
Published: (2025)
by: Gao, Jialin, et al.
Published: (2025)
Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors
by: Feng, Yingjie, et al.
Published: (2026)
by: Feng, Yingjie, et al.
Published: (2026)
SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
by: Wang, Yinqiao, et al.
Published: (2024)
by: Wang, Yinqiao, et al.
Published: (2024)
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
Medical Large Vision Language Models with Multi-Image Visual Ability
by: Yang, Xikai, et al.
Published: (2025)
by: Yang, Xikai, et al.
Published: (2025)
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
by: Jin, Xin, et al.
Published: (2025)
by: Jin, Xin, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
by: XU, Dunyuan, et al.
Published: (2025)
by: XU, Dunyuan, et al.
Published: (2025)
S^2Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
by: Pei, Jialun, et al.
Published: (2024)
by: Pei, Jialun, et al.
Published: (2024)
Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor Segmentation
by: Yang, Kaixiang, et al.
Published: (2024)
by: Yang, Kaixiang, et al.
Published: (2024)
UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data
by: Zou, Longkun, et al.
Published: (2025)
by: Zou, Longkun, et al.
Published: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Towards Synchronous Memorizability and Generalizability with Site-Modulated Diffusion Replay for Cross-Site Continual Segmentation
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
by: Wang, Jinghao, et al.
Published: (2026)
by: Wang, Jinghao, et al.
Published: (2026)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
by: Qi, Yu, et al.
Published: (2026)
by: Qi, Yu, et al.
Published: (2026)
SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation
by: Chen, Weihong, et al.
Published: (2025)
by: Chen, Weihong, et al.
Published: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
by: Liu, Yichen, et al.
Published: (2026)
by: Liu, Yichen, et al.
Published: (2026)
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
by: Wang, Yinqiao, et al.
Published: (2025)
by: Wang, Yinqiao, et al.
Published: (2025)
Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
by: Goswami, Dipam, et al.
Published: (2026)
by: Goswami, Dipam, et al.
Published: (2026)
Adaptive Negative Evidential Deep Learning for Open-set Semi-supervised Learning
by: Yu, Yang, et al.
Published: (2023)
by: Yu, Yang, et al.
Published: (2023)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
by: Yu, Hai-Tao, et al.
Published: (2024)
by: Yu, Hai-Tao, et al.
Published: (2024)
MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation
by: Cai, Kaixin, et al.
Published: (2023)
by: Cai, Kaixin, et al.
Published: (2023)
Similar Items
-
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024) -
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024) -
MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
by: Zhou, Donghao, et al.
Published: (2024) -
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025) -
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)