Gespeichert in:
| Hauptverfasser: | Zhan, Yu-Wei, Wu, Xiao-Ming, Luo, Xin, Wei, Yinwei, Xu, Xin-Shun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2406.10776 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Deep Mamba Multi-modal Learning
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
von: Yuan, Xu, et al.
Veröffentlicht: (2023)
von: Yuan, Xu, et al.
Veröffentlicht: (2023)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
von: He, Xusheng, et al.
Veröffentlicht: (2025)
von: He, Xusheng, et al.
Veröffentlicht: (2025)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
How to Cache Important Contents for Multi-modal Service in Dynamic Networks: A DRL-based Caching Scheme
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
von: Fu, Jingwen, et al.
Veröffentlicht: (2025)
von: Fu, Jingwen, et al.
Veröffentlicht: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
von: Li, Zhenyang, et al.
Veröffentlicht: (2023)
von: Li, Zhenyang, et al.
Veröffentlicht: (2023)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
von: An, Xiao, et al.
Veröffentlicht: (2026)
von: An, Xiao, et al.
Veröffentlicht: (2026)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
von: Li, Bingliang, et al.
Veröffentlicht: (2024)
von: Li, Bingliang, et al.
Veröffentlicht: (2024)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
MDF: A Dynamic Fusion Model for Multi-modal Fake News Detection
von: Lv, Hongzhen, et al.
Veröffentlicht: (2024)
von: Lv, Hongzhen, et al.
Veröffentlicht: (2024)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
von: Zhao, Xianbing, et al.
Veröffentlicht: (2025)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2025)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
von: Wu, Shilong
Veröffentlicht: (2025)
von: Wu, Shilong
Veröffentlicht: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2024) -
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025) -
Deep Mamba Multi-modal Learning
von: Zhu, Jian, et al.
Veröffentlicht: (2024) -
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025) -
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
von: Jiang, Xin, et al.
Veröffentlicht: (2025)