Enhancing CLIP Robustness via Cross-Modality Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Xingyu, Zhu, Beier, Wang, Shuo, Zhao, Kesen, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Semantic Alignment for Image Clustering
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
by: Zhao, Kesen, et al.
Published: (2025)
by: Zhao, Kesen, et al.
Published: (2025)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
by: Zhao, Kesen, et al.
Published: (2026)
by: Zhao, Kesen, et al.
Published: (2026)
Selective Vision-Language Subspace Projection for Few-shot CLIP
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis
by: Han, Yaohui, et al.
Published: (2026)
by: Han, Yaohui, et al.
Published: (2026)
CollaMamba: Efficient Collaborative Perception with Cross-Agent Spatial-Temporal State Space Model
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning
by: Qiu, Yilun, et al.
Published: (2026)
by: Qiu, Yilun, et al.
Published: (2026)
Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance
by: Wang, Xiangxiang, et al.
Published: (2025)
by: Wang, Xiangxiang, et al.
Published: (2025)
V2X-DGPE: Addressing Domain Gaps and Pose Errors for Robust Collaborative 3D Object Detection
by: Wang, Sichao, et al.
Published: (2025)
by: Wang, Sichao, et al.
Published: (2025)
TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception
by: Song, Zhiying, et al.
Published: (2025)
by: Song, Zhiying, et al.
Published: (2025)
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Robust Fine-tuning of Zero-shot Models via Variance Reduction
by: Zhu, Beier, et al.
Published: (2024)
by: Zhu, Beier, et al.
Published: (2024)
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
ReCCur: A Recursive Corner-Case Curation Framework for Robust Vision-Language Understanding in Open and Edge Scenarios
by: Wei, Yihan, et al.
Published: (2026)
by: Wei, Yihan, et al.
Published: (2026)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
PixelFlowCast: Latent-Free Precipitation Nowcasting via Pixel Mean Flows
by: Zhu, Yufeng, et al.
Published: (2026)
by: Zhu, Yufeng, et al.
Published: (2026)
Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG
by: Madavan, Rakesh Raj, et al.
Published: (2025)
by: Madavan, Rakesh Raj, et al.
Published: (2025)
Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learning
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
by: Wang, Zirui, et al.
Published: (2024)
by: Wang, Zirui, et al.
Published: (2024)
How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI
by: Wen, Bo, et al.
Published: (2025)
by: Wen, Bo, et al.
Published: (2025)
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
by: Shi, Haoyuan, et al.
Published: (2025)
by: Shi, Haoyuan, et al.
Published: (2025)
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
by: Chung-En, et al.
Published: (2025)
by: Chung-En, et al.
Published: (2025)
Fast2comm:Collaborative perception combined with prior knowledge
by: Zhang, Zhengbin, et al.
Published: (2025)
by: Zhang, Zhengbin, et al.
Published: (2025)
Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification
by: Rehman, Tayyab, et al.
Published: (2026)
by: Rehman, Tayyab, et al.
Published: (2026)
AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
by: Chen, Jiao, et al.
Published: (2025)
by: Chen, Jiao, et al.
Published: (2025)
SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
by: Tu, Rong-Cheng, et al.
Published: (2024)
by: Tu, Rong-Cheng, et al.
Published: (2024)
PreGSU-A Generalized Traffic Scene Understanding Model for Autonomous Driving based on Pre-trained Graph Attention Network
by: Wang, Yuning, et al.
Published: (2024)
by: Wang, Yuning, et al.
Published: (2024)
Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling
by: Yu, Chung-En Johnny, et al.
Published: (2025)
by: Yu, Chung-En Johnny, et al.
Published: (2025)
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
by: Li, Yuankai, et al.
Published: (2026)
by: Li, Yuankai, et al.
Published: (2026)
AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
by: Fathi, Nima, et al.
Published: (2025)
by: Fathi, Nima, et al.
Published: (2025)
Allen: Rethinking MAS Design through Step-Level Policy Autonomy
by: Zhou, Qiangong, et al.
Published: (2025)
by: Zhou, Qiangong, et al.
Published: (2025)
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
by: Lin, Xiangye, et al.
Published: (2026)
by: Lin, Xiangye, et al.
Published: (2026)
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization
by: Meng, Chutian, et al.
Published: (2026)
by: Meng, Chutian, et al.
Published: (2026)
ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection
by: Xu, Yingjia, et al.
Published: (2026)
by: Xu, Yingjia, et al.
Published: (2026)
SPACE: 3D Spatial Co-operation and Exploration Framework for Robust Mapping and Coverage with Multi-Robot Systems
by: Ghanta, Sai Krishna, et al.
Published: (2024)
by: Ghanta, Sai Krishna, et al.
Published: (2024)
Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance
by: Fan, Hongxing, et al.
Published: (2025)
by: Fan, Hongxing, et al.
Published: (2025)
Unified End-to-End V2X Cooperative Autonomous Driving
by: Li, Zhiwei, et al.
Published: (2024)
by: Li, Zhiwei, et al.
Published: (2024)
Similar Items
-
Hierarchical Semantic Alignment for Image Clustering
by: Zhu, Xingyu, et al.
Published: (2025) -
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
by: Zhu, Xingyu, et al.
Published: (2026) -
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
by: Zhao, Kesen, et al.
Published: (2025) -
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025) -
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)