Saved in:
| Main Authors: | Qin, Yang, Chen, Chao, Fu, Zhihang, Peng, Dezhong, Peng, Xi, Hu, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.11036 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification
by: Li, Yongxiang, et al.
Published: (2025)
by: Li, Yongxiang, et al.
Published: (2025)
Robust Fuzzy Multi-view Learning under View Conflict
by: Duan, Siyuan, et al.
Published: (2026)
by: Duan, Siyuan, et al.
Published: (2026)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Prototypical Prompting for Text-to-image Person Re-identification
by: Yan, Shuanglin, et al.
Published: (2024)
by: Yan, Shuanglin, et al.
Published: (2024)
Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
by: Wang, Shijie, et al.
Published: (2025)
by: Wang, Shijie, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Deep Reversible Consistency Learning for Cross-modal Retrieval
by: Pu, Ruitao, et al.
Published: (2025)
by: Pu, Ruitao, et al.
Published: (2025)
Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
by: Chao, Qin, et al.
Published: (2025)
by: Chao, Qin, et al.
Published: (2025)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
by: Munyer, Travis, et al.
Published: (2023)
by: Munyer, Travis, et al.
Published: (2023)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)
by: Zheng, Guan, et al.
Published: (2025)
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
by: Zhang, Yafei, et al.
Published: (2025)
by: Zhang, Yafei, et al.
Published: (2025)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
by: Wu, Ruiqi, et al.
Published: (2024)
by: Wu, Ruiqi, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Deep Learning-based Text-in-Image Watermarking
by: Karki, Bishwa, et al.
Published: (2024)
by: Karki, Bishwa, et al.
Published: (2024)
Hybrid Feedback-Guided Optimal Learning for Wireless Interactive Panoramic Scene Delivery
by: Wu, Xiaoyi, et al.
Published: (2026)
by: Wu, Xiaoyi, et al.
Published: (2026)
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
by: Ahtesham, Muhammad, et al.
Published: (2025)
by: Ahtesham, Muhammad, et al.
Published: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
VDCook:DIY video data cook your MLLMs
by: Wu, Chengwei
Published: (2026)
by: Wu, Chengwei
Published: (2026)
Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention
by: Liu, Moyang, et al.
Published: (2025)
by: Liu, Moyang, et al.
Published: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
CAMERA: Adapting to Semantic Camouflage in Unsupervised Text-Attributed Graph Fraud Detection
by: Pan, Junjun, et al.
Published: (2026)
by: Pan, Junjun, et al.
Published: (2026)
The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation
by: Bertazzini, Giulia, et al.
Published: (2025)
by: Bertazzini, Giulia, et al.
Published: (2025)
Enhancing Cross-Prompt Transferability in Vision-Language Models through Contextual Injection of Target Tokens
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
AI-Integrated Decision Support System for Real-Time Market Growth Forecasting and Multi-Source Content Diffusion Analytics
by: Yin, Ziqing, et al.
Published: (2025)
by: Yin, Ziqing, et al.
Published: (2025)
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
by: Hoe, Jiun Tian, et al.
Published: (2023)
by: Hoe, Jiun Tian, et al.
Published: (2023)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
by: Wang, Juncheng, et al.
Published: (2025)
by: Wang, Juncheng, et al.
Published: (2025)
Demonstration of MaskSearch: Efficiently Querying Image Masks for Machine Learning Workflows
by: Wei, Lindsey Linxi, et al.
Published: (2024)
by: Wei, Lindsey Linxi, et al.
Published: (2024)
Stable Multimodal Graph Unlearning via Feature-Dimension Aware Quantile Selection
by: Zhou, Jingjing, et al.
Published: (2026)
by: Zhou, Jingjing, et al.
Published: (2026)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
Exploring Modality Disruption in Multimodal Fake News Detection
by: Liu, Moyang, et al.
Published: (2025)
by: Liu, Moyang, et al.
Published: (2025)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
Multimodal Representation Learning and Fusion
by: Jin, Qihang, et al.
Published: (2025)
by: Jin, Qihang, et al.
Published: (2025)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
by: Pu, Ruitao, et al.
Published: (2025)
by: Pu, Ruitao, et al.
Published: (2025)
Similar Items
-
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023) -
Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification
by: Li, Yongxiang, et al.
Published: (2025) -
Robust Fuzzy Multi-view Learning under View Conflict
by: Duan, Siyuan, et al.
Published: (2026) -
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025) -
Prototypical Prompting for Text-to-image Person Re-identification
by: Yan, Shuanglin, et al.
Published: (2024)