Guardado en:
| Autores principales: | Yan, Shuanglin, Liu, Jun, Dong, Neng, Zhang, Liyan, Tang, Jinhui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.09427 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Embedding and Enriching Explicit Semantics for Visible-Infrared Person Re-Identification
por: Dong, Neng, et al.
Publicado: (2024)
por: Dong, Neng, et al.
Publicado: (2024)
Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-Identification
por: Dong, Neng, et al.
Publicado: (2025)
por: Dong, Neng, et al.
Publicado: (2025)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
por: Qin, Yang, et al.
Publicado: (2023)
por: Qin, Yang, et al.
Publicado: (2023)
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
por: Qin, Yang, et al.
Publicado: (2025)
por: Qin, Yang, et al.
Publicado: (2025)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
por: Wu, Ruiqi, et al.
Publicado: (2024)
por: Wu, Ruiqi, et al.
Publicado: (2024)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
por: Tang, Hao, et al.
Publicado: (2026)
por: Tang, Hao, et al.
Publicado: (2026)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
por: Shu, Ying, et al.
Publicado: (2026)
por: Shu, Ying, et al.
Publicado: (2026)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
por: Cui, Can, et al.
Publicado: (2024)
por: Cui, Can, et al.
Publicado: (2024)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition
por: Zhang, Zhicheng, et al.
Publicado: (2025)
por: Zhang, Zhicheng, et al.
Publicado: (2025)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
Learning Shared Sentiment Prototypes for Adaptive Multimodal Sentiment Analysis
por: Su, Chen, et al.
Publicado: (2026)
por: Su, Chen, et al.
Publicado: (2026)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
por: Wang, Hanyao, et al.
Publicado: (2024)
por: Wang, Hanyao, et al.
Publicado: (2024)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
por: Jiang, Xin, et al.
Publicado: (2024)
por: Jiang, Xin, et al.
Publicado: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
por: He, Jiayi, et al.
Publicado: (2025)
por: He, Jiayi, et al.
Publicado: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
por: Xie, Jingjing, et al.
Publicado: (2024)
por: Xie, Jingjing, et al.
Publicado: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
por: Wang, Yongqi, et al.
Publicado: (2025)
por: Wang, Yongqi, et al.
Publicado: (2025)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
por: Zhang, Deyu, et al.
Publicado: (2025)
por: Zhang, Deyu, et al.
Publicado: (2025)
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
por: Ma, Zhiyong, et al.
Publicado: (2025)
por: Ma, Zhiyong, et al.
Publicado: (2025)
Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
por: Zhao, Xianbing, et al.
Publicado: (2025)
por: Zhao, Xianbing, et al.
Publicado: (2025)
Personalized Image Generation with Large Multimodal Models
por: Xu, Yiyan, et al.
Publicado: (2024)
por: Xu, Yiyan, et al.
Publicado: (2024)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
por: Niu, Xinlei, et al.
Publicado: (2024)
por: Niu, Xinlei, et al.
Publicado: (2024)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
por: Xie, Zequn, et al.
Publicado: (2026)
por: Xie, Zequn, et al.
Publicado: (2026)
Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification
por: Li, Yongxiang, et al.
Publicado: (2025)
por: Li, Yongxiang, et al.
Publicado: (2025)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
por: Hei, Nailei, et al.
Publicado: (2024)
por: Hei, Nailei, et al.
Publicado: (2024)
DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
por: Chi, Kaiyi, et al.
Publicado: (2025)
por: Chi, Kaiyi, et al.
Publicado: (2025)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
por: Wang, Yuhao, et al.
Publicado: (2025)
por: Wang, Yuhao, et al.
Publicado: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
por: Quan, Weize, et al.
Publicado: (2024)
por: Quan, Weize, et al.
Publicado: (2024)
Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
por: Ziatdinov, Rushan, et al.
Publicado: (2025)
por: Ziatdinov, Rushan, et al.
Publicado: (2025)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
por: Jiang, Chaoya, et al.
Publicado: (2023)
por: Jiang, Chaoya, et al.
Publicado: (2023)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
por: Qiu, Shi, et al.
Publicado: (2026)
por: Qiu, Shi, et al.
Publicado: (2026)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
por: Gao, Jiayi, et al.
Publicado: (2025)
por: Gao, Jiayi, et al.
Publicado: (2025)
ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
por: Zhang, Haoshuo, et al.
Publicado: (2025)
por: Zhang, Haoshuo, et al.
Publicado: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
por: Yang, Shuyu, et al.
Publicado: (2024)
por: Yang, Shuyu, et al.
Publicado: (2024)
CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval
por: Qin, Yawen, et al.
Publicado: (2026)
por: Qin, Yawen, et al.
Publicado: (2026)
Video Streaming with Kairos: An MPC-Based ABR with Streaming-Aware Throughput Prediction
por: Zhong, Ziyu, et al.
Publicado: (2025)
por: Zhong, Ziyu, et al.
Publicado: (2025)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
por: Gui, Yinxuan, et al.
Publicado: (2025)
por: Gui, Yinxuan, et al.
Publicado: (2025)
Learning Switchable Priors for Neural Image Compression
por: Zhang, Haotian, et al.
Publicado: (2025)
por: Zhang, Haotian, et al.
Publicado: (2025)
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
Ejemplares similares
-
Embedding and Enriching Explicit Semantics for Visible-Infrared Person Re-Identification
por: Dong, Neng, et al.
Publicado: (2024) -
Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-Identification
por: Dong, Neng, et al.
Publicado: (2025) -
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
por: Qin, Yang, et al.
Publicado: (2023) -
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
por: Qin, Yang, et al.
Publicado: (2025) -
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
por: Wu, Ruiqi, et al.
Publicado: (2024)