Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jinlong, Zhao, Dong, Jie, Zequn, Ricci, Elisa, Ma, Lin, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2024)
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
von: Zhao, Dong, et al.
Veröffentlicht: (2025)
von: Zhao, Dong, et al.
Veröffentlicht: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
von: Li, Jinlong, et al.
Veröffentlicht: (2026)
von: Li, Jinlong, et al.
Veröffentlicht: (2026)
Democratizing Fine-grained Visual Recognition with Large Language Models
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
Large-scale Pre-trained Models are Surprisingly Strong in Incremental Novel Class Discovery
von: Liu, Mingxuan, et al.
Veröffentlicht: (2023)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2023)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
von: Xing, Songlong, et al.
Veröffentlicht: (2026)
von: Xing, Songlong, et al.
Veröffentlicht: (2026)
Vision+X: A Survey on Multimodal Learning in the Light of Data
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
Large Language Models for Multimodal Deformable Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
von: Garcia-Fernandez, Pablo, et al.
Veröffentlicht: (2025)
von: Garcia-Fernandez, Pablo, et al.
Veröffentlicht: (2025)
LESS: Label-Efficient and Single-Stage Referring 3D Segmentation
von: Liu, Xuexun, et al.
Veröffentlicht: (2024)
von: Liu, Xuexun, et al.
Veröffentlicht: (2024)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
von: Lv, Chonghua, et al.
Veröffentlicht: (2026)
von: Lv, Chonghua, et al.
Veröffentlicht: (2026)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
von: Zuo, Zhi, et al.
Veröffentlicht: (2025)
von: Zuo, Zhi, et al.
Veröffentlicht: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
von: Ge, Xuri, et al.
Veröffentlicht: (2024)
von: Ge, Xuri, et al.
Veröffentlicht: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
von: Betti, Federico, et al.
Veröffentlicht: (2024)
von: Betti, Federico, et al.
Veröffentlicht: (2024)
Group Orthogonalization Regularization For Vision Models Adaptation and Robustness
von: Kurtz, Yoav, et al.
Veröffentlicht: (2023)
von: Kurtz, Yoav, et al.
Veröffentlicht: (2023)
Safe Vision-Language Models via Unsafe Weights Manipulation
von: D'Incà, Moreno, et al.
Veröffentlicht: (2025)
von: D'Incà, Moreno, et al.
Veröffentlicht: (2025)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
von: Ye, Weixin, et al.
Veröffentlicht: (2026)
von: Ye, Weixin, et al.
Veröffentlicht: (2026)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2025)
Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
von: Zhao, Dong, et al.
Veröffentlicht: (2026)
von: Zhao, Dong, et al.
Veröffentlicht: (2026)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning
von: Huang, Duojun, et al.
Veröffentlicht: (2024)
von: Huang, Duojun, et al.
Veröffentlicht: (2024)
Reverse Personalization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
Asymmetric GANs for Image-to-Image Translation
von: Tang, Hao, et al.
Veröffentlicht: (2019)
von: Tang, Hao, et al.
Veröffentlicht: (2019)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
von: Chen, Shaoxiang, et al.
Veröffentlicht: (2024)
von: Chen, Shaoxiang, et al.
Veröffentlicht: (2024)
Stable Neighbor Denoising for Source-free Domain Adaptive Segmentation
von: Zhao, Dong, et al.
Veröffentlicht: (2024)
von: Zhao, Dong, et al.
Veröffentlicht: (2024)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
Causal Disentanglement for Robust Long-tail Medical Image Generation
von: Nie, Weizhi, et al.
Veröffentlicht: (2025)
von: Nie, Weizhi, et al.
Veröffentlicht: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
von: Song, Yue, et al.
Veröffentlicht: (2023)
von: Song, Yue, et al.
Veröffentlicht: (2023)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
von: Li, Jinlong, et al.
Veröffentlicht: (2025) -
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2024) -
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
von: Zhao, Dong, et al.
Veröffentlicht: (2025) -
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
von: Li, Jinlong, et al.
Veröffentlicht: (2026) -
Democratizing Fine-grained Visual Recognition with Large Language Models
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)