Vision+X: A Survey on Multimodal Learning in the Light of Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Ye, Wu, Yu, Sebe, Nicu, Yan, Yan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
por: Shu, Yan, et al.
Publicado: (2025)
por: Shu, Yan, et al.
Publicado: (2025)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
por: Li, Jinlong, et al.
Publicado: (2024)
por: Li, Jinlong, et al.
Publicado: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
por: Wang, Weijie, et al.
Publicado: (2022)
por: Wang, Weijie, et al.
Publicado: (2022)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
por: Xing, Songlong, et al.
Publicado: (2025)
por: Xing, Songlong, et al.
Publicado: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
por: Zhang, Ji, et al.
Publicado: (2025)
por: Zhang, Ji, et al.
Publicado: (2025)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
por: Shu, Yan, et al.
Publicado: (2025)
por: Shu, Yan, et al.
Publicado: (2025)
Hierarchical Cross-Attention Network for Virtual Try-On
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
por: Ge, Xuri, et al.
Publicado: (2024)
por: Ge, Xuri, et al.
Publicado: (2024)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
por: Garcia-Fernandez, Pablo, et al.
Publicado: (2025)
por: Garcia-Fernandez, Pablo, et al.
Publicado: (2025)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
por: Ye, Weixin, et al.
Publicado: (2026)
por: Ye, Weixin, et al.
Publicado: (2026)
Reverse Personalization
por: Kung, Han-Wei, et al.
Publicado: (2025)
por: Kung, Han-Wei, et al.
Publicado: (2025)
Large Language Models for Multimodal Deformable Image Registration
por: Ma, Mingrui, et al.
Publicado: (2024)
por: Ma, Mingrui, et al.
Publicado: (2024)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
por: Zheng, Xu, et al.
Publicado: (2025)
por: Zheng, Xu, et al.
Publicado: (2025)
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
por: Zhao, Dong, et al.
Publicado: (2025)
por: Zhao, Dong, et al.
Publicado: (2025)
Asymmetric GANs for Image-to-Image Translation
por: Tang, Hao, et al.
Publicado: (2019)
por: Tang, Hao, et al.
Publicado: (2019)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
por: Xing, Songlong, et al.
Publicado: (2026)
por: Xing, Songlong, et al.
Publicado: (2026)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
por: Song, Yue, et al.
Publicado: (2023)
por: Song, Yue, et al.
Publicado: (2023)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
por: Li, Jinlong, et al.
Publicado: (2025)
por: Li, Jinlong, et al.
Publicado: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
por: Li, Jinlong, et al.
Publicado: (2026)
por: Li, Jinlong, et al.
Publicado: (2026)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
por: Liu, Jiaqi, et al.
Publicado: (2025)
por: Liu, Jiaqi, et al.
Publicado: (2025)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
por: Shu, Yan, et al.
Publicado: (2026)
por: Shu, Yan, et al.
Publicado: (2026)
Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
por: Liu, Jian, et al.
Publicado: (2024)
por: Liu, Jian, et al.
Publicado: (2024)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
por: Lv, Chonghua, et al.
Publicado: (2026)
por: Lv, Chonghua, et al.
Publicado: (2026)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
por: Zuo, Zhi, et al.
Publicado: (2025)
por: Zuo, Zhi, et al.
Publicado: (2025)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
NullFace: Training-Free Localized Face Anonymization
por: Kung, Han-Wei, et al.
Publicado: (2025)
por: Kung, Han-Wei, et al.
Publicado: (2025)
Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
por: Zheng, Haiyang, et al.
Publicado: (2025)
por: Zheng, Haiyang, et al.
Publicado: (2025)
Masked Image Modeling: A Survey
por: Hondru, Vlad, et al.
Publicado: (2024)
por: Hondru, Vlad, et al.
Publicado: (2024)
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
por: He, Yikang, et al.
Publicado: (2026)
por: He, Yikang, et al.
Publicado: (2026)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
por: Xue, Feng, et al.
Publicado: (2025)
por: Xue, Feng, et al.
Publicado: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
por: Peruzzo, Elia, et al.
Publicado: (2025)
por: Peruzzo, Elia, et al.
Publicado: (2025)
Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
por: Zheng, Haiyang, et al.
Publicado: (2024)
por: Zheng, Haiyang, et al.
Publicado: (2024)
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
por: Zheng, Haiyang, et al.
Publicado: (2025)
por: Zheng, Haiyang, et al.
Publicado: (2025)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
por: Betti, Federico, et al.
Publicado: (2024)
por: Betti, Federico, et al.
Publicado: (2024)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
por: Wang, Yidan, et al.
Publicado: (2025)
por: Wang, Yidan, et al.
Publicado: (2025)
Hallucination Early Detection in Diffusion Models
por: Betti, Federico, et al.
Publicado: (2026)
por: Betti, Federico, et al.
Publicado: (2026)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
por: Zheng, Haiyang, et al.
Publicado: (2024)
por: Zheng, Haiyang, et al.
Publicado: (2024)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
por: Sacilotti, André, et al.
Publicado: (2024)
por: Sacilotti, André, et al.
Publicado: (2024)
Hyperbolic Busemann Neural Networks
por: Chen, Ziheng, et al.
Publicado: (2026)
por: Chen, Ziheng, et al.
Publicado: (2026)
Ejemplares similares
-
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
por: Shu, Yan, et al.
Publicado: (2025) -
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
por: Li, Jinlong, et al.
Publicado: (2024) -
Rethinking the Learning Paradigm for Facial Expression Recognition
por: Wang, Weijie, et al.
Publicado: (2022) -
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
por: Xing, Songlong, et al.
Publicado: (2025) -
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
por: Zhang, Ji, et al.
Publicado: (2025)