Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yi, Zhang, Ce |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Variational Adapter for Cross-modal Similarity Representation
von: Wei, WenZhang, et al.
Veröffentlicht: (2026)
von: Wei, WenZhang, et al.
Veröffentlicht: (2026)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
von: Cai, Rui, et al.
Veröffentlicht: (2024)
von: Cai, Rui, et al.
Veröffentlicht: (2024)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
von: Hu, Yunzuo, et al.
Veröffentlicht: (2026)
von: Hu, Yunzuo, et al.
Veröffentlicht: (2026)
Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
von: Gong, Yan, et al.
Veröffentlicht: (2025)
von: Gong, Yan, et al.
Veröffentlicht: (2025)
FunOTTA: On-the-Fly Adaptation on Cross-Domain Fundus Image via Stable Test-time Training
von: Zeng, Qian, et al.
Veröffentlicht: (2024)
von: Zeng, Qian, et al.
Veröffentlicht: (2024)
EVCtrl: Efficient Control Adapter for Visual Generation
von: Yang, Zixiang, et al.
Veröffentlicht: (2025)
von: Yang, Zixiang, et al.
Veröffentlicht: (2025)
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
von: Dong, Sibo, et al.
Veröffentlicht: (2025)
von: Dong, Sibo, et al.
Veröffentlicht: (2025)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
von: Zhang, Tianyu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyu, et al.
Veröffentlicht: (2025)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
Learning Visual Conditioning Tokens to Correct Domain Shift for Fully Test-time Adaptation
von: Tang, Yushun, et al.
Veröffentlicht: (2024)
von: Tang, Yushun, et al.
Veröffentlicht: (2024)
Test-time Alignment-Enhanced Adapter for Vision-Language Models
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
von: Zhang, Yurong, et al.
Veröffentlicht: (2024)
von: Zhang, Yurong, et al.
Veröffentlicht: (2024)
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
von: Guo, Meng-Hao, et al.
Veröffentlicht: (2025)
von: Guo, Meng-Hao, et al.
Veröffentlicht: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
von: Xing, Ling, et al.
Veröffentlicht: (2024)
von: Xing, Ling, et al.
Veröffentlicht: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
von: Zhang, Runmin, et al.
Veröffentlicht: (2024)
von: Zhang, Runmin, et al.
Veröffentlicht: (2024)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
von: Tao, Mingzhe, et al.
Veröffentlicht: (2026)
von: Tao, Mingzhe, et al.
Veröffentlicht: (2026)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
von: Wang, Hu, et al.
Veröffentlicht: (2023)
von: Wang, Hu, et al.
Veröffentlicht: (2023)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
Concept-Guided Prompt Learning for Generalization in Vision-Language Models
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
von: Chen, Junan, et al.
Veröffentlicht: (2025)
von: Chen, Junan, et al.
Veröffentlicht: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Learning Triangular Distribution in Visual World
von: Chen, Ping, et al.
Veröffentlicht: (2023)
von: Chen, Ping, et al.
Veröffentlicht: (2023)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
von: Lu, Wenting, et al.
Veröffentlicht: (2026)
von: Lu, Wenting, et al.
Veröffentlicht: (2026)
ΩSFormer: Dual-Modal Ω-like Super-Resolution Transformer Network for Cross-scale and High-accuracy Terraced Field Vectorization Extraction
von: Li, Chang, et al.
Veröffentlicht: (2024)
von: Li, Chang, et al.
Veröffentlicht: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
von: Gou, Yunhao, et al.
Veröffentlicht: (2025)
von: Gou, Yunhao, et al.
Veröffentlicht: (2025)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
Rethinking Unsupervised Cross-modal Flow Estimation: Learning from Decoupled Optimization and Consistency Constraint
von: Zhang, Runmin, et al.
Veröffentlicht: (2025)
von: Zhang, Runmin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Variational Adapter for Cross-modal Similarity Representation
von: Wei, WenZhang, et al.
Veröffentlicht: (2026) -
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
von: Cai, Rui, et al.
Veröffentlicht: (2024) -
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023) -
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
von: Hu, Yunzuo, et al.
Veröffentlicht: (2026) -
Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)