Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Shiyao, Wang, Zi-An, Yin, Kangning, Tian, Zheng, Zhang, Mingyuan, Si, Weixin, Zou, Shihao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tri-Modal Motion Retrieval by Learning a Joint Embedding Space
di: Yin, Kangning, et al.
Pubblicazione: (2024)
di: Yin, Kangning, et al.
Pubblicazione: (2024)
Semantics-Aware Human Motion Generation from Audio Instructions
di: Wang, Zi-An, et al.
Pubblicazione: (2025)
di: Wang, Zi-An, et al.
Pubblicazione: (2025)
RACon: Retrieval-Augmented Simulated Character Locomotion Control
di: Mu, Yuxuan, et al.
Pubblicazione: (2024)
di: Mu, Yuxuan, et al.
Pubblicazione: (2024)
Highly Efficient 3D Human Pose Tracking from Events with Spiking Spatiotemporal Transformer
di: Zou, Shihao, et al.
Pubblicazione: (2023)
di: Zou, Shihao, et al.
Pubblicazione: (2023)
TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular Registration
di: Liu, Zehua, et al.
Pubblicazione: (2026)
di: Liu, Zehua, et al.
Pubblicazione: (2026)
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
di: Li, Shihao, et al.
Pubblicazione: (2025)
di: Li, Shihao, et al.
Pubblicazione: (2025)
Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
di: Galoaa, Bishoy, et al.
Pubblicazione: (2025)
di: Galoaa, Bishoy, et al.
Pubblicazione: (2025)
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
di: Hossain, Md Aminur, et al.
Pubblicazione: (2026)
di: Hossain, Md Aminur, et al.
Pubblicazione: (2026)
Large Motion Model for Unified Multi-Modal Motion Generation
di: Zhang, Mingyuan, et al.
Pubblicazione: (2024)
di: Zhang, Mingyuan, et al.
Pubblicazione: (2024)
WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval
di: Ren, Junlong, et al.
Pubblicazione: (2025)
di: Ren, Junlong, et al.
Pubblicazione: (2025)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
di: Xue, Junxiao, et al.
Pubblicazione: (2026)
di: Xue, Junxiao, et al.
Pubblicazione: (2026)
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
di: Zhang, Yao, et al.
Pubblicazione: (2026)
di: Zhang, Yao, et al.
Pubblicazione: (2026)
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
di: Chen, Hanmo, et al.
Pubblicazione: (2026)
di: Chen, Hanmo, et al.
Pubblicazione: (2026)
Multi-Grained Compositional Visual Clue Learning for Image Intent Recognition
di: Tang, Yin, et al.
Pubblicazione: (2025)
di: Tang, Yin, et al.
Pubblicazione: (2025)
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
di: Wang, Shijie, et al.
Pubblicazione: (2026)
di: Wang, Shijie, et al.
Pubblicazione: (2026)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
di: Wang, Yiqun, et al.
Pubblicazione: (2024)
di: Wang, Yiqun, et al.
Pubblicazione: (2024)
Toward Real-Time Surgical Scene Segmentation via a Spike-Driven Video Transformer with Spike-Informed Pretraining
di: Zou, Shihao, et al.
Pubblicazione: (2025)
di: Zou, Shihao, et al.
Pubblicazione: (2025)
Learning Parallax for Stereo Event-based Motion Deblurring
di: Lin, Mingyuan, et al.
Pubblicazione: (2023)
di: Lin, Mingyuan, et al.
Pubblicazione: (2023)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
di: Wang, Wei, et al.
Pubblicazione: (2024)
di: Wang, Wei, et al.
Pubblicazione: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
di: Zhou, Junjie, et al.
Pubblicazione: (2024)
di: Zhou, Junjie, et al.
Pubblicazione: (2024)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
di: Walmer, Matthew, et al.
Pubblicazione: (2023)
di: Walmer, Matthew, et al.
Pubblicazione: (2023)
Language-driven Fine-grained Retrieval
di: Wang, Shijie, et al.
Pubblicazione: (2025)
di: Wang, Shijie, et al.
Pubblicazione: (2025)
Multi-Modal Generative Embedding Model
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
di: Liang, Weixin, et al.
Pubblicazione: (2025)
di: Liang, Weixin, et al.
Pubblicazione: (2025)
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
di: Zhu, Minghao, et al.
Pubblicazione: (2023)
di: Zhu, Minghao, et al.
Pubblicazione: (2023)
Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding Approach
di: Zou, Mian, et al.
Pubblicazione: (2024)
di: Zou, Mian, et al.
Pubblicazione: (2024)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
di: Fang, Haopeng, et al.
Pubblicazione: (2024)
di: Fang, Haopeng, et al.
Pubblicazione: (2024)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
di: Du, Yipeng, et al.
Pubblicazione: (2025)
di: Du, Yipeng, et al.
Pubblicazione: (2025)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
di: Shen, Keming, et al.
Pubblicazione: (2025)
di: Shen, Keming, et al.
Pubblicazione: (2025)
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
di: Trinh, Quoc-Huy
Pubblicazione: (2025)
di: Trinh, Quoc-Huy
Pubblicazione: (2025)
MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning
di: Ma, Hongxu, et al.
Pubblicazione: (2025)
di: Ma, Hongxu, et al.
Pubblicazione: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
di: Zhang, Zhengxuan, et al.
Pubblicazione: (2025)
di: Zhang, Zhengxuan, et al.
Pubblicazione: (2025)
PAS-Mamba: Phase-Amplitude-Spatial State Space Model for MRI Reconstruction
di: Kui, Xiaoyan, et al.
Pubblicazione: (2026)
di: Kui, Xiaoyan, et al.
Pubblicazione: (2026)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
di: Korbar, Bruno, et al.
Pubblicazione: (2025)
di: Korbar, Bruno, et al.
Pubblicazione: (2025)
Retrieval Robust to Object Motion Blur
di: Zou, Rong, et al.
Pubblicazione: (2024)
di: Zou, Rong, et al.
Pubblicazione: (2024)
Modality-Agnostic Structural Image Representation Learning for Deformable Multi-Modality Medical Image Registration
di: Mok, Tony C. W., et al.
Pubblicazione: (2024)
di: Mok, Tony C. W., et al.
Pubblicazione: (2024)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
di: Fei, Zhengcong, et al.
Pubblicazione: (2023)
di: Fei, Zhengcong, et al.
Pubblicazione: (2023)
Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic Learning
di: Tang, Haomiao, et al.
Pubblicazione: (2026)
di: Tang, Haomiao, et al.
Pubblicazione: (2026)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
di: Huang, Jincai, et al.
Pubblicazione: (2026)
di: Huang, Jincai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Tri-Modal Motion Retrieval by Learning a Joint Embedding Space
di: Yin, Kangning, et al.
Pubblicazione: (2024) -
Semantics-Aware Human Motion Generation from Audio Instructions
di: Wang, Zi-An, et al.
Pubblicazione: (2025) -
RACon: Retrieval-Augmented Simulated Character Locomotion Control
di: Mu, Yuxuan, et al.
Pubblicazione: (2024) -
Highly Efficient 3D Human Pose Tracking from Events with Spiking Spatiotemporal Transformer
di: Zou, Shihao, et al.
Pubblicazione: (2023) -
TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular Registration
di: Liu, Zehua, et al.
Pubblicazione: (2026)