Robust Domain Generalization for Multi-modal Object Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Qiao, Yuxin, Li, Keqin, Lin, Junhong, Wei, Rong, Jiang, Chufeng, Luo, Yang, Yang, Haoyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games
by: Yi, Hanling, et al.
Published: (2026)
by: Yi, Hanling, et al.
Published: (2026)
CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
by: Liu, Lemin, et al.
Published: (2025)
by: Liu, Lemin, et al.
Published: (2025)
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
by: Yang, Xudong, et al.
Published: (2025)
by: Yang, Xudong, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
by: Wu, Linquan, et al.
Published: (2026)
by: Wu, Linquan, et al.
Published: (2026)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
Domain Similarity-Perceived Label Assignment for Domain Generalized Underwater Object Detection
by: Li, Xisheng, et al.
Published: (2023)
by: Li, Xisheng, et al.
Published: (2023)
Interact-Custom: Customized Human Object Interaction Image Generation
by: Xu, Zhu, et al.
Published: (2025)
by: Xu, Zhu, et al.
Published: (2025)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024)
by: Zhao, Lirui, et al.
Published: (2024)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
MapFusion: A Novel BEV Feature Fusion Network for Multi-modal Map Construction
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
Human Activity Recognition using RGB-Event based Sensors: A Multi-modal Heat Conduction Model and A Benchmark Dataset
by: Wang, Shiao, et al.
Published: (2025)
by: Wang, Shiao, et al.
Published: (2025)
CD-FKD: Cross-Domain Feature Knowledge Distillation for Robust Single-Domain Generalization in Object Detection
by: Lee, Junseok, et al.
Published: (2026)
by: Lee, Junseok, et al.
Published: (2026)
PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection
by: Xu, Xiaoran, et al.
Published: (2024)
by: Xu, Xiaoran, et al.
Published: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Detection
by: Liu, Jiacong, et al.
Published: (2026)
by: Liu, Jiacong, et al.
Published: (2026)
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
by: Zuo, Lin, et al.
Published: (2024)
by: Zuo, Lin, et al.
Published: (2024)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
by: Baby, Britty, et al.
Published: (2025)
by: Baby, Britty, et al.
Published: (2025)
Vision-Language Consistency Guided Multi-modal Prompt Learning for Blind AI Generated Image Quality Assessment
by: Fu, Jun, et al.
Published: (2024)
by: Fu, Jun, et al.
Published: (2024)
Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving
by: Lou, Yang, et al.
Published: (2023)
by: Lou, Yang, et al.
Published: (2023)
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024)
by: Wei, Shicai, et al.
Published: (2024)
MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment
by: Liu, Tianyi, et al.
Published: (2024)
by: Liu, Tianyi, et al.
Published: (2024)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Generative RLHF-V: Learning Principles from Multi-modal Human Preference
by: Zhou, Jiayi, et al.
Published: (2025)
by: Zhou, Jiayi, et al.
Published: (2025)
Graph-Attention Network with Adversarial Domain Alignment for Robust Cross-Domain Facial Expression Recognition
by: Ghaedi, Razieh, et al.
Published: (2025)
by: Ghaedi, Razieh, et al.
Published: (2025)
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
by: Shangguan, Zeyu, et al.
Published: (2025)
by: Shangguan, Zeyu, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
by: Yoon, Heegeon, et al.
Published: (2026)
by: Yoon, Heegeon, et al.
Published: (2026)
Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
by: Xuan, Shiyu, et al.
Published: (2023)
by: Xuan, Shiyu, et al.
Published: (2023)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
by: Dai, Ming, et al.
Published: (2024)
by: Dai, Ming, et al.
Published: (2024)
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
by: Jiang, Yutao, et al.
Published: (2025)
by: Jiang, Yutao, et al.
Published: (2025)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
by: Liu, Shenxi, et al.
Published: (2025)
by: Liu, Shenxi, et al.
Published: (2025)
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
by: Zhu, Armando, et al.
Published: (2024)
by: Zhu, Armando, et al.
Published: (2024)
Similar Items
-
Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games
by: Yi, Hanling, et al.
Published: (2026) -
CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
by: Liu, Lemin, et al.
Published: (2025) -
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
by: Yang, Xudong, et al.
Published: (2025) -
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024) -
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
by: He, Hulingxiao, et al.
Published: (2026)