Continual Cross-Modal Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Yan, Huang, Hai, Fang, Minghui, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts
by: Zhao, Shijia, et al.
Published: (2025)
by: Zhao, Shijia, et al.
Published: (2025)
GCC: Generative Calibration Clustering
by: Xia, Haifeng, et al.
Published: (2024)
by: Xia, Haifeng, et al.
Published: (2024)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions
by: Chong, Jun-Xiong, et al.
Published: (2025)
by: Chong, Jun-Xiong, et al.
Published: (2025)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
DiffX: Guide Your Layout to Cross-Modal Generative Modeling
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
by: Liang, Yongyuan, et al.
Published: (2025)
by: Liang, Yongyuan, et al.
Published: (2025)
Cross-Modal Clinical Knowledge Integration for Mammography Report Generation
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet
by: Niu, Xin, et al.
Published: (2026)
by: Niu, Xin, et al.
Published: (2026)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
Dual-Level Cross-Modal Contrastive Clustering
by: Zhang, Haixin, et al.
Published: (2024)
by: Zhang, Haixin, et al.
Published: (2024)
Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
by: Wu, Chen, et al.
Published: (2026)
by: Wu, Chen, et al.
Published: (2026)
MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification
by: Zhao, Yujian, et al.
Published: (2025)
by: Zhao, Yujian, et al.
Published: (2025)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
Ultrasound Report Generation with Cross-Modality Feature Alignment via Unsupervised Guidance
by: Li, Jun, et al.
Published: (2024)
by: Li, Jun, et al.
Published: (2024)
Federated Cross-Modal Style-Aware Prompt Generation
by: Prasad, Suraj, et al.
Published: (2025)
by: Prasad, Suraj, et al.
Published: (2025)
Cross-Modal Causal Intervention for Medical Report Generation
by: Chen, Weixing, et al.
Published: (2023)
by: Chen, Weixing, et al.
Published: (2023)
BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment
by: Zhou, Kanglei, et al.
Published: (2026)
by: Zhou, Kanglei, et al.
Published: (2026)
BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
by: Wen, Haiquan, et al.
Published: (2025)
by: Wen, Haiquan, et al.
Published: (2025)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
by: Huang, Kuan Wei, et al.
Published: (2025)
by: Huang, Kuan Wei, et al.
Published: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
by: Zhou, Hefeng, et al.
Published: (2026)
by: Zhou, Hefeng, et al.
Published: (2026)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
by: Luo, Yuqing, et al.
Published: (2025)
by: Luo, Yuqing, et al.
Published: (2025)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
by: Huang, Shunyu, et al.
Published: (2026)
by: Huang, Shunyu, et al.
Published: (2026)
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Cross-Modality Gait Recognition: Bridging LiDAR and Camera Modalities for Human Identification
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
by: Wang, Hang, et al.
Published: (2026)
by: Wang, Hang, et al.
Published: (2026)
Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
by: Liu, Weide, et al.
Published: (2025)
by: Liu, Weide, et al.
Published: (2025)
Cross-Modal Consistency Learning for Sign Language Recognition
by: Wu, Kepeng, et al.
Published: (2025)
by: Wu, Kepeng, et al.
Published: (2025)
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
by: Du, Yuxuan, et al.
Published: (2025)
by: Du, Yuxuan, et al.
Published: (2025)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
Multi-Modal Generative Embedding Model
by: Ma, Feipeng, et al.
Published: (2024)
by: Ma, Feipeng, et al.
Published: (2024)
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
User-Friendly Customized Generation with Multi-Modal Prompts
by: Zhong, Linhao, et al.
Published: (2024)
by: Zhong, Linhao, et al.
Published: (2024)
Similar Items
-
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025) -
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024) -
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025) -
SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts
by: Zhao, Shijia, et al.
Published: (2025) -
GCC: Generative Calibration Clustering
by: Xia, Haifeng, et al.
Published: (2024)