GeoMM: On Geodesic Perspective for Multi-modal Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mei, Shibin, Wang, Hang, Ni, Bingbing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MSSIDD: A Benchmark for Multi-Sensor Denoising
von: Mei, Shibin, et al.
Veröffentlicht: (2024)
von: Mei, Shibin, et al.
Veröffentlicht: (2024)
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
von: Li, Leheng, et al.
Veröffentlicht: (2024)
von: Li, Leheng, et al.
Veröffentlicht: (2024)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
von: Xuan, Shiyu, et al.
Veröffentlicht: (2025)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2025)
From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning
von: Ni, Hang, et al.
Veröffentlicht: (2026)
von: Ni, Hang, et al.
Veröffentlicht: (2026)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
von: Wu, Chenyang, et al.
Veröffentlicht: (2024)
von: Wu, Chenyang, et al.
Veröffentlicht: (2024)
A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
von: Wang, Xiaodong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaodong, et al.
Veröffentlicht: (2025)
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
von: Deichler, Anna, et al.
Veröffentlicht: (2024)
von: Deichler, Anna, et al.
Veröffentlicht: (2024)
Multi-modality action recognition based on dual feature shift in vehicle cabin monitoring
von: Lin, Dan, et al.
Veröffentlicht: (2024)
von: Lin, Dan, et al.
Veröffentlicht: (2024)
GeoDiffMM: Geometry-Guided Conditional Diffusion for Motion Magnification
von: Liu, Xuedeng, et al.
Veröffentlicht: (2025)
von: Liu, Xuedeng, et al.
Veröffentlicht: (2025)
MM-GTUNets: Unified Multi-Modal Graph Deep Learning for Brain Disorders Prediction
von: Cai, Luhui, et al.
Veröffentlicht: (2024)
von: Cai, Luhui, et al.
Veröffentlicht: (2024)
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
von: Cao, Yuhang, et al.
Veröffentlicht: (2024)
von: Cao, Yuhang, et al.
Veröffentlicht: (2024)
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
von: Dong, Zichao, et al.
Veröffentlicht: (2024)
von: Dong, Zichao, et al.
Veröffentlicht: (2024)
GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis
von: Wang, Xuqin, et al.
Veröffentlicht: (2026)
von: Wang, Xuqin, et al.
Veröffentlicht: (2026)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
PM2: A New Prompting Multi-modal Model Paradigm for Few-shot Medical Image Classification
von: Wang, Zhenwei, et al.
Veröffentlicht: (2024)
von: Wang, Zhenwei, et al.
Veröffentlicht: (2024)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
von: Jiang, Yangzhou, et al.
Veröffentlicht: (2024)
von: Jiang, Yangzhou, et al.
Veröffentlicht: (2024)
GeoCD: A Differential Local Approximation for Geodesic Chamfer Distance
von: Alonso, Pedro, et al.
Veröffentlicht: (2025)
von: Alonso, Pedro, et al.
Veröffentlicht: (2025)
GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
von: Chen, Yuedong, et al.
Veröffentlicht: (2020)
von: Chen, Yuedong, et al.
Veröffentlicht: (2020)
MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers
von: Dong, Zichao, et al.
Veröffentlicht: (2024)
von: Dong, Zichao, et al.
Veröffentlicht: (2024)
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
von: Li, Yuanze, et al.
Veröffentlicht: (2024)
von: Li, Yuanze, et al.
Veröffentlicht: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
Human-AI Collaborative Multi-modal Multi-rater Learning for Endometriosis Diagnosis
von: Wang, Hu, et al.
Veröffentlicht: (2024)
von: Wang, Hu, et al.
Veröffentlicht: (2024)
SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes
von: Zohaib, Mohammad, et al.
Veröffentlicht: (2024)
von: Zohaib, Mohammad, et al.
Veröffentlicht: (2024)
DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation
von: Xiong, Yuxuan, et al.
Veröffentlicht: (2025)
von: Xiong, Yuxuan, et al.
Veröffentlicht: (2025)
3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
von: Zhu, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhu, Yupeng, et al.
Veröffentlicht: (2025)
PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
von: Jin, Qiaoqiao, et al.
Veröffentlicht: (2024)
von: Jin, Qiaoqiao, et al.
Veröffentlicht: (2024)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
von: Sun, Mingzhen, et al.
Veröffentlicht: (2024)
von: Sun, Mingzhen, et al.
Veröffentlicht: (2024)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
von: Wang, Teng, et al.
Veröffentlicht: (2024)
von: Wang, Teng, et al.
Veröffentlicht: (2024)
Vectorized Video Representation with Easy Editing via Hierarchical Spatio-Temporally Consistent Proxy Embedding
von: Chen, Ye, et al.
Veröffentlicht: (2025)
von: Chen, Ye, et al.
Veröffentlicht: (2025)
TAL: Two-stream Adaptive Learning for Generalizable Person Re-identification
von: Yan, Yichao, et al.
Veröffentlicht: (2021)
von: Yan, Yichao, et al.
Veröffentlicht: (2021)
Learning Geodesics of Geometric Shape Deformations From Images
von: Wu, Nian, et al.
Veröffentlicht: (2024)
von: Wu, Nian, et al.
Veröffentlicht: (2024)
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
von: Wang, Kunpeng, et al.
Veröffentlicht: (2024)
von: Wang, Kunpeng, et al.
Veröffentlicht: (2024)
HumanMM: Global Human Motion Recovery from Multi-shot Videos
von: Zhang, Yuhong, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MSSIDD: A Benchmark for Multi-Sensor Denoising
von: Mei, Shibin, et al.
Veröffentlicht: (2024) -
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025) -
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
von: Wen, Jiahao, et al.
Veröffentlicht: (2025) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024) -
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
von: Li, Leheng, et al.
Veröffentlicht: (2024)