Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zeyu, Zhang, Pengfei, Ye, Kai, Dong, Wei, Feng, Xin, Zhang, Yana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026)
by: He, Yuxuan, et al.
Published: (2026)
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
by: Gomi, Keisuke, et al.
Published: (2026)
by: Gomi, Keisuke, et al.
Published: (2026)
Understanding Identity Continuity in Thermal Video through Scene-Level Consistency
by: Sun, Wei-Chieh, et al.
Published: (2026)
by: Sun, Wei-Chieh, et al.
Published: (2026)
Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition
by: Chen, Kuan Wei, et al.
Published: (2026)
by: Chen, Kuan Wei, et al.
Published: (2026)
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
by: Li, Lin, et al.
Published: (2026)
by: Li, Lin, et al.
Published: (2026)
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
Improving Visual Object Tracking through Visual Prompting
by: Chen, Shih-Fang, et al.
Published: (2024)
by: Chen, Shih-Fang, et al.
Published: (2024)
Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal Retrieval
by: Liu, Yizhi, et al.
Published: (2025)
by: Liu, Yizhi, et al.
Published: (2025)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
by: Zhang, Junbin, et al.
Published: (2026)
by: Zhang, Junbin, et al.
Published: (2026)
Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal
by: Leng, Yicheng, et al.
Published: (2025)
by: Leng, Yicheng, et al.
Published: (2025)
Do Inpainting Yourself: Generative Facial Inpainting Guided by Exemplars
by: Lu, Wanglong, et al.
Published: (2022)
by: Lu, Wanglong, et al.
Published: (2022)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
Prevailing Research Areas for Music AI in the Era of Foundation Models
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
MetaErr: Towards Predicting Error Patterns in Deep Neural Networks
by: Totakura, Varun, et al.
Published: (2026)
by: Totakura, Varun, et al.
Published: (2026)
Digital analysis of early color photographs taken using regular color screen processes
by: Hubička, Jan, et al.
Published: (2023)
by: Hubička, Jan, et al.
Published: (2023)
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
by: Zhang, Pengfei, et al.
Published: (2026)
by: Zhang, Pengfei, et al.
Published: (2026)
Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models
by: Tralie, Christopher J., et al.
Published: (2024)
by: Tralie, Christopher J., et al.
Published: (2024)
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
by: Korolkov, Vasilii, et al.
Published: (2025)
by: Korolkov, Vasilii, et al.
Published: (2025)
Motion Attribution for Video Generation
by: Wu, Xindi, et al.
Published: (2026)
by: Wu, Xindi, et al.
Published: (2026)
ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection
by: Samson, Hema Hariharan
Published: (2026)
by: Samson, Hema Hariharan
Published: (2026)
M3LEO: A Multi-Modal, Multi-Label Earth Observation Dataset Integrating Interferometric SAR and Multispectral Data
by: Allen, Matthew J, et al.
Published: (2024)
by: Allen, Matthew J, et al.
Published: (2024)
Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Pinching Visuo-haptic Display: Investigating Cross-Modal Effects of Visual Textures on Electrostatic Cloth Tactile Sensations
by: Kitagishi, Takekazu, et al.
Published: (2025)
by: Kitagishi, Takekazu, et al.
Published: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
by: Choi, Seonghwa, et al.
Published: (2025)
by: Choi, Seonghwa, et al.
Published: (2025)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
GOT-JEPA: Generic Object Tracking with Model Adaptation and Occlusion Handling using Joint-Embedding Predictive Architecture
by: Chen, Shih-Fang, et al.
Published: (2026)
by: Chen, Shih-Fang, et al.
Published: (2026)
GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model Editing
by: Chen, Shih-Fang, et al.
Published: (2026)
by: Chen, Shih-Fang, et al.
Published: (2026)
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
by: Lee, Seungmi, et al.
Published: (2025)
by: Lee, Seungmi, et al.
Published: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Lightweight Complementary-Cue Fusion for Robust Video Face Forgery Detection
by: Baek, Sunghwan, et al.
Published: (2026)
by: Baek, Sunghwan, et al.
Published: (2026)
AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models
by: Yun, Kwan, et al.
Published: (2025)
by: Yun, Kwan, et al.
Published: (2025)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
by: Shu, Hongzhi, et al.
Published: (2024)
by: Shu, Hongzhi, et al.
Published: (2024)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
by: Li, Shuyu, et al.
Published: (2025)
by: Li, Shuyu, et al.
Published: (2025)
Mixture of Experts Approaches in Dense Retrieval Tasks
by: Sokli, Effrosyni, et al.
Published: (2025)
by: Sokli, Effrosyni, et al.
Published: (2025)
Saliency-Aware Diffusion Reconstruction for Effective Invisible Watermark Removal
by: Alam, Inzamamul, et al.
Published: (2025)
by: Alam, Inzamamul, et al.
Published: (2025)
Multi-level SSL Feature Gating for Audio Deepfake Detection
by: Tran, Hoan My, et al.
Published: (2025)
by: Tran, Hoan My, et al.
Published: (2025)
Similar Items
-
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026) -
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
by: Gomi, Keisuke, et al.
Published: (2026) -
Understanding Identity Continuity in Thermal Video through Scene-Level Consistency
by: Sun, Wei-Chieh, et al.
Published: (2026) -
Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition
by: Chen, Kuan Wei, et al.
Published: (2026) -
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
by: Li, Lin, et al.
Published: (2026)