Salvato in:
| Autori principali: | Wang, Xiao, Wang, Shiao, Ding, Yuhe, Li, Yuehang, Wu, Wentao, Rong, Yao, Kong, Weizhe, Huang, Ju, Li, Shihao, Yang, Haoxiang, Wang, Ziwen, Jiang, Bo, Li, Chenglong, Wang, Yaowei, Tian, Yonghong, Tang, Jin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2404.09516 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
di: Wang, Xiao, et al.
Pubblicazione: (2023)
di: Wang, Xiao, et al.
Pubblicazione: (2023)
SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
di: Wang, Xiao, et al.
Pubblicazione: (2023)
di: Wang, Xiao, et al.
Pubblicazione: (2023)
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
di: Wang, Xiao, et al.
Pubblicazione: (2026)
di: Wang, Xiao, et al.
Pubblicazione: (2026)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
di: Yuan, Bo, et al.
Pubblicazione: (2024)
di: Yuan, Bo, et al.
Pubblicazione: (2024)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
di: Lu, Zhenyu, et al.
Pubblicazione: (2025)
di: Lu, Zhenyu, et al.
Pubblicazione: (2025)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose Estimation
di: Fu, Yajie, et al.
Pubblicazione: (2025)
di: Fu, Yajie, et al.
Pubblicazione: (2025)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
di: Lian, Niu, et al.
Pubblicazione: (2025)
di: Lian, Niu, et al.
Pubblicazione: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024)
di: Zhang, Bo, et al.
Pubblicazione: (2024)
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm
di: Jin, Jiandong, et al.
Pubblicazione: (2023)
di: Jin, Jiandong, et al.
Pubblicazione: (2023)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
di: Li, Yaowei, et al.
Pubblicazione: (2025)
di: Li, Yaowei, et al.
Pubblicazione: (2025)
Order Is Not Layout: Order-to-Space Bias in Image Generation
di: Zhang, Yongkang, et al.
Pubblicazione: (2026)
di: Zhang, Yongkang, et al.
Pubblicazione: (2026)
Image Conductor: Precision Control for Interactive Video Synthesis
di: Li, Yaowei, et al.
Pubblicazione: (2024)
di: Li, Yaowei, et al.
Pubblicazione: (2024)
Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production
di: Feng, Liqian, et al.
Pubblicazione: (2025)
di: Feng, Liqian, et al.
Pubblicazione: (2025)
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
di: Ouyang, Shuyi, et al.
Pubblicazione: (2024)
di: Ouyang, Shuyi, et al.
Pubblicazione: (2024)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
di: Wang, Bing, et al.
Pubblicazione: (2025)
di: Wang, Bing, et al.
Pubblicazione: (2025)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
di: Li, Liupeng, et al.
Pubblicazione: (2026)
di: Li, Liupeng, et al.
Pubblicazione: (2026)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
di: Wang, Song, et al.
Pubblicazione: (2025)
di: Wang, Song, et al.
Pubblicazione: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
di: Li, Jun, et al.
Pubblicazione: (2025)
di: Li, Jun, et al.
Pubblicazione: (2025)
HR-INR: Continuous Space-Time Video Super-Resolution via Event Camera
di: Lu, Yunfan, et al.
Pubblicazione: (2024)
di: Lu, Yunfan, et al.
Pubblicazione: (2024)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
di: Liu, Lingyu, et al.
Pubblicazione: (2026)
di: Liu, Lingyu, et al.
Pubblicazione: (2026)
Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
di: Nie, Fan, et al.
Pubblicazione: (2025)
di: Nie, Fan, et al.
Pubblicazione: (2025)
Deep Shape-Texture Statistics for Completely Blind Image Quality Evaluation
di: Li, Yixuan, et al.
Pubblicazione: (2024)
di: Li, Yixuan, et al.
Pubblicazione: (2024)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
di: Wang, Bing, et al.
Pubblicazione: (2025)
di: Wang, Bing, et al.
Pubblicazione: (2025)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
di: Lu, Zhenyu, et al.
Pubblicazione: (2026)
di: Lu, Zhenyu, et al.
Pubblicazione: (2026)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
di: Wang, Bing, et al.
Pubblicazione: (2024)
di: Wang, Bing, et al.
Pubblicazione: (2024)
Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
A Survey of Information Disorder on Video-Sharing Platforms
di: Li, Meiyu, et al.
Pubblicazione: (2025)
di: Li, Meiyu, et al.
Pubblicazione: (2025)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
di: Kong, Chenqi, et al.
Pubblicazione: (2023)
di: Kong, Chenqi, et al.
Pubblicazione: (2023)
CartoAgent: a multimodal large language model-powered multi-agent cartographic framework for map style transfer and evaluation
di: Wang, Chenglong, et al.
Pubblicazione: (2025)
di: Wang, Chenglong, et al.
Pubblicazione: (2025)
L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression
di: Zhang, Junxuan, et al.
Pubblicazione: (2024)
di: Zhang, Junxuan, et al.
Pubblicazione: (2024)
Retrieval-Augmented Multimodal Model for Fake News Detection
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
di: Luo, Anwei, et al.
Pubblicazione: (2023)
di: Luo, Anwei, et al.
Pubblicazione: (2023)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
A Novel Approach to Industrial Defect Generation through Blended Latent Diffusion Model with Online Adaptation
di: Li, Hanxi, et al.
Pubblicazione: (2024)
di: Li, Hanxi, et al.
Pubblicazione: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
di: Zhang, Zhenxing, et al.
Pubblicazione: (2024)
di: Zhang, Zhenxing, et al.
Pubblicazione: (2024)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
di: Zhang, Wang, et al.
Pubblicazione: (2024)
di: Zhang, Wang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
di: Wang, Xiao, et al.
Pubblicazione: (2023) -
SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
di: Wang, Xiao, et al.
Pubblicazione: (2023) -
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
di: Wang, Xiao, et al.
Pubblicazione: (2026) -
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
di: Yuan, Bo, et al.
Pubblicazione: (2024) -
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
di: Lu, Zhenyu, et al.
Pubblicazione: (2025)