Gespeichert in:
| Hauptverfasser: | Jia, Haonan, Dong, Shichao, Dong, Xin, Sun, Zenghui, Wang, Jin, Lan, Jinsong, Zhu, Xiaoyong, Zheng, Bo, Zhang, Kaifu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.01696 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
von: Dong, Xin, et al.
Veröffentlicht: (2025)
von: Dong, Xin, et al.
Veröffentlicht: (2025)
REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
von: Zheng, Jun, et al.
Veröffentlicht: (2026)
von: Zheng, Jun, et al.
Veröffentlicht: (2026)
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework
von: Ling, Zihan, et al.
Veröffentlicht: (2025)
von: Ling, Zihan, et al.
Veröffentlicht: (2025)
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
von: Wang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Wang, Zhicheng, et al.
Veröffentlicht: (2025)
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
von: Song, Quanjian, et al.
Veröffentlicht: (2026)
von: Song, Quanjian, et al.
Veröffentlicht: (2026)
Enhancing Sequential Recommendation with World Knowledge from Large Language Models
von: Dai, Tianjie, et al.
Veröffentlicht: (2025)
von: Dai, Tianjie, et al.
Veröffentlicht: (2025)
Towards RGB-NIR Cross-modality Image Registration and Beyond
von: Li, Huadong, et al.
Veröffentlicht: (2024)
von: Li, Huadong, et al.
Veröffentlicht: (2024)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
Improving Container Port Efficiency: A Data‐Driven Model for Optimizing Truck Arrival Appointments Through Distributionally Robust Optimization
von: Shichao Sun, et al.
Veröffentlicht: (2025)
von: Shichao Sun, et al.
Veröffentlicht: (2025)
Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment
von: Song, Jia, et al.
Veröffentlicht: (2024)
von: Song, Jia, et al.
Veröffentlicht: (2024)
Self Identity Mapping
von: Cai, Xiuding, et al.
Veröffentlicht: (2025)
von: Cai, Xiuding, et al.
Veröffentlicht: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
Leveraging Scene Context with Dual Networks for Sequential User Behavior Modeling
von: Chen, Xu, et al.
Veröffentlicht: (2025)
von: Chen, Xu, et al.
Veröffentlicht: (2025)
Cross-Modal Distillation For Widely Differing Modalities
von: Zhao, Cairong, et al.
Veröffentlicht: (2025)
von: Zhao, Cairong, et al.
Veröffentlicht: (2025)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
von: Chen, Lei, et al.
Veröffentlicht: (2026)
von: Chen, Lei, et al.
Veröffentlicht: (2026)
Information and communications technologies for carbon sinks from economics and engineering perspectives
von: Dong, Yuze, et al.
Veröffentlicht: (2026)
von: Dong, Yuze, et al.
Veröffentlicht: (2026)
Mobile UMI: Cross-View Diffusion Policy with Decoupled Kinematics for Mobile Manipulation
von: Huang, Haoran, et al.
Veröffentlicht: (2026)
von: Huang, Haoran, et al.
Veröffentlicht: (2026)
Learning Multi-Branch Cooperation for Enhanced Click-Through Rate Prediction at Taobao
von: Chen, Xu, et al.
Veröffentlicht: (2024)
von: Chen, Xu, et al.
Veröffentlicht: (2024)
AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear Mapping
von: Dong, Haonan, et al.
Veröffentlicht: (2025)
von: Dong, Haonan, et al.
Veröffentlicht: (2025)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
von: Cai, Rui, et al.
Veröffentlicht: (2024)
von: Cai, Rui, et al.
Veröffentlicht: (2024)
Networked Restless Multi-Arm Bandits with Reinforcement Learning
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
Network Embedding via Deep Prediction Model
von: Sun, Xin, et al.
Veröffentlicht: (2021)
von: Sun, Xin, et al.
Veröffentlicht: (2021)
ICFNet: Integrated Cross-modal Fusion Network for Survival Prediction
von: Zhang, Binyu, et al.
Veröffentlicht: (2025)
von: Zhang, Binyu, et al.
Veröffentlicht: (2025)
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
von: Yue, Zhengrong, et al.
Veröffentlicht: (2026)
von: Yue, Zhengrong, et al.
Veröffentlicht: (2026)
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
von: Zheng, Baihui, et al.
Veröffentlicht: (2025)
von: Zheng, Baihui, et al.
Veröffentlicht: (2025)
A Review of Detection, Evolution, and Data Reconstruction Strategies for False Data Injection Attacks in Power Cyber-Physical Systems
von: Bo, Xiaoyong
Veröffentlicht: (2025)
von: Bo, Xiaoyong
Veröffentlicht: (2025)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
Cellulose‐Reinforced Bilayer Hydrogel Actuator With Solvent‐Responsiveness
von: Zenghui Li, et al.
Veröffentlicht: (2026)
von: Zenghui Li, et al.
Veröffentlicht: (2026)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
von: Chen, Yizhu, et al.
Veröffentlicht: (2025)
von: Chen, Yizhu, et al.
Veröffentlicht: (2025)
Rethinking Adam for Time Series Forecasting: A Simple Heuristic to Improve Optimization under Distribution Shifts
von: Dong, Yuze, et al.
Veröffentlicht: (2026)
von: Dong, Yuze, et al.
Veröffentlicht: (2026)
Quasinormal modes of charged covariant effective black holes with a cosmological constant
von: Dong, Zhongzhinan, et al.
Veröffentlicht: (2026)
von: Dong, Zhongzhinan, et al.
Veröffentlicht: (2026)
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
von: Liu, Tao, et al.
Veröffentlicht: (2026)
von: Liu, Tao, et al.
Veröffentlicht: (2026)
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges
von: Shen, Huangjun, et al.
Veröffentlicht: (2024)
von: Shen, Huangjun, et al.
Veröffentlicht: (2024)
Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
Electrothermally‐Driven Carbon Nanotube Fiber Artificial Muscle With Endomysium‐Inspired Sheath and Multifilament Core
von: Sufeng Zhu, et al.
Veröffentlicht: (2025)
von: Sufeng Zhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
von: Dong, Xin, et al.
Veröffentlicht: (2025) -
REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
von: Tang, Yiwen, et al.
Veröffentlicht: (2025) -
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
von: Zheng, Jun, et al.
Veröffentlicht: (2026) -
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework
von: Ling, Zihan, et al.
Veröffentlicht: (2025) -
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
von: Wang, Zhicheng, et al.
Veröffentlicht: (2025)