GeoMM: On Geodesic Perspective for Multi-modal Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Mei, Shibin, Wang, Hang, Ni, Bingbing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MSSIDD: A Benchmark for Multi-Sensor Denoising
by: Mei, Shibin, et al.
Published: (2024)
by: Mei, Shibin, et al.
Published: (2024)
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
by: Wen, Jiahao, et al.
Published: (2025)
by: Wen, Jiahao, et al.
Published: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
by: Xuan, Shiyu, et al.
Published: (2025)
by: Xuan, Shiyu, et al.
Published: (2025)
From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning
by: Ni, Hang, et al.
Published: (2026)
by: Ni, Hang, et al.
Published: (2026)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
by: Wu, Chenyang, et al.
Published: (2024)
by: Wu, Chenyang, et al.
Published: (2024)
A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
by: Deichler, Anna, et al.
Published: (2024)
by: Deichler, Anna, et al.
Published: (2024)
Multi-modality action recognition based on dual feature shift in vehicle cabin monitoring
by: Lin, Dan, et al.
Published: (2024)
by: Lin, Dan, et al.
Published: (2024)
GeoDiffMM: Geometry-Guided Conditional Diffusion for Motion Magnification
by: Liu, Xuedeng, et al.
Published: (2025)
by: Liu, Xuedeng, et al.
Published: (2025)
MM-GTUNets: Unified Multi-Modal Graph Deep Learning for Brain Disorders Prediction
by: Cai, Luhui, et al.
Published: (2024)
by: Cai, Luhui, et al.
Published: (2024)
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
by: Cao, Yuhang, et al.
Published: (2024)
by: Cao, Yuhang, et al.
Published: (2024)
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
by: Dong, Zichao, et al.
Published: (2024)
by: Dong, Zichao, et al.
Published: (2024)
GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis
by: Wang, Xuqin, et al.
Published: (2026)
by: Wang, Xuqin, et al.
Published: (2026)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
PM2: A New Prompting Multi-modal Model Paradigm for Few-shot Medical Image Classification
by: Wang, Zhenwei, et al.
Published: (2024)
by: Wang, Zhenwei, et al.
Published: (2024)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
by: Jiang, Yangzhou, et al.
Published: (2024)
by: Jiang, Yangzhou, et al.
Published: (2024)
GeoCD: A Differential Local Approximation for Geodesic Chamfer Distance
by: Alonso, Pedro, et al.
Published: (2025)
by: Alonso, Pedro, et al.
Published: (2025)
GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
by: Chen, Yuedong, et al.
Published: (2020)
by: Chen, Yuedong, et al.
Published: (2020)
MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers
by: Dong, Zichao, et al.
Published: (2024)
by: Dong, Zichao, et al.
Published: (2024)
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
by: Li, Yuanze, et al.
Published: (2024)
by: Li, Yuanze, et al.
Published: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
by: Xu, Tianyu, et al.
Published: (2025)
by: Xu, Tianyu, et al.
Published: (2025)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Human-AI Collaborative Multi-modal Multi-rater Learning for Endometriosis Diagnosis
by: Wang, Hu, et al.
Published: (2024)
by: Wang, Hu, et al.
Published: (2024)
SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes
by: Zohaib, Mohammad, et al.
Published: (2024)
by: Zohaib, Mohammad, et al.
Published: (2024)
DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation
by: Xiong, Yuxuan, et al.
Published: (2025)
by: Xiong, Yuxuan, et al.
Published: (2025)
3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
by: Zhu, Yupeng, et al.
Published: (2025)
by: Zhu, Yupeng, et al.
Published: (2025)
PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation
by: Hu, Zhangli, et al.
Published: (2025)
by: Hu, Zhangli, et al.
Published: (2025)
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
by: Jin, Qiaoqiao, et al.
Published: (2024)
by: Jin, Qiaoqiao, et al.
Published: (2024)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Vectorized Video Representation with Easy Editing via Hierarchical Spatio-Temporally Consistent Proxy Embedding
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
TAL: Two-stream Adaptive Learning for Generalizable Person Re-identification
by: Yan, Yichao, et al.
Published: (2021)
by: Yan, Yichao, et al.
Published: (2021)
Learning Geodesics of Geometric Shape Deformations From Images
by: Wu, Nian, et al.
Published: (2024)
by: Wu, Nian, et al.
Published: (2024)
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
by: Wang, Kunpeng, et al.
Published: (2024)
by: Wang, Kunpeng, et al.
Published: (2024)
HumanMM: Global Human Motion Recovery from Multi-shot Videos
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Similar Items
-
MSSIDD: A Benchmark for Multi-Sensor Denoising
by: Mei, Shibin, et al.
Published: (2024) -
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
by: Wang, Jiaqi, et al.
Published: (2025) -
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
by: Wen, Jiahao, et al.
Published: (2025) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024) -
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024)