Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ruan, Shouwei, Dong, Yinpeng, Liu, Hanqing, Huang, Yao, Su, Hang, Wei, Xingxing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?
by: Ruan, Shouwei, et al.
Published: (2024)
by: Ruan, Shouwei, et al.
Published: (2024)
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
by: Liu, Hanqing, et al.
Published: (2025)
by: Liu, Hanqing, et al.
Published: (2025)
Towards Transferable Targeted 3D Adversarial Attack in the Physical World
by: Huang, Yao, et al.
Published: (2023)
by: Huang, Yao, et al.
Published: (2023)
Real-world Adversarial Defense against Patch Attacks based on Diffusion Model
by: Wei, Xingxing, et al.
Published: (2024)
by: Wei, Xingxing, et al.
Published: (2024)
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
by: Kang, Caixin, et al.
Published: (2023)
by: Kang, Caixin, et al.
Published: (2023)
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
by: Huang, Zhuo, et al.
Published: (2023)
by: Huang, Zhuo, et al.
Published: (2023)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
by: Ruan, Shouwei, et al.
Published: (2025)
by: Ruan, Shouwei, et al.
Published: (2025)
NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
by: Sun, Yitong, et al.
Published: (2025)
by: Sun, Yitong, et al.
Published: (2025)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
FaceCat: Enhancing Face Recognition Security with a Unified Diffusion Model
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Improving Viewpoint-Invariance and Temporal Consistency for Action Detection
by: Porto, Yannick, et al.
Published: (2026)
by: Porto, Yannick, et al.
Published: (2026)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
by: Zhang, Peng-Fei, et al.
Published: (2025)
by: Zhang, Peng-Fei, et al.
Published: (2025)
Efficient and Effective Universal Adversarial Attack against Vision-Language Pre-training Models
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
VILA: On Pre-training for Visual Language Models
by: Lin, Ji, et al.
Published: (2023)
by: Lin, Ji, et al.
Published: (2023)
Rethinking Model Ensemble in Transfer-based Adversarial Attacks
by: Chen, Huanran, et al.
Published: (2023)
by: Chen, Huanran, et al.
Published: (2023)
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
by: Zhang, Duzhen, et al.
Published: (2025)
by: Zhang, Duzhen, et al.
Published: (2025)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
SDPT: Synchronous Dual Prompt Tuning for Fusion-based Visual-Language Pre-trained Models
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
Boosting Image Restoration via Priors from Pre-trained Models
by: Xu, Xiaogang, et al.
Published: (2024)
by: Xu, Xiaogang, et al.
Published: (2024)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
3D Scene Graph Guided Vision-Language Pre-training
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
by: Zheng, Haonan, et al.
Published: (2024)
by: Zheng, Haonan, et al.
Published: (2024)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training
by: Wang, Rongsheng, et al.
Published: (2023)
by: Wang, Rongsheng, et al.
Published: (2023)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
by: Tian, Yuxin, et al.
Published: (2024)
by: Tian, Yuxin, et al.
Published: (2024)
Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
by: Yang, Yiting, et al.
Published: (2025)
by: Yang, Yiting, et al.
Published: (2025)
One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Continual Forgetting for Pre-trained Vision Models
by: Zhao, Hongbo, et al.
Published: (2024)
by: Zhao, Hongbo, et al.
Published: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Split Adaptation for Pre-trained Vision Transformers
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
Similar Items
-
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?
by: Ruan, Shouwei, et al.
Published: (2024) -
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
by: Liu, Hanqing, et al.
Published: (2025) -
Towards Transferable Targeted 3D Adversarial Attack in the Physical World
by: Huang, Yao, et al.
Published: (2023) -
Real-world Adversarial Defense against Patch Attacks based on Diffusion Model
by: Wei, Xingxing, et al.
Published: (2024) -
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
by: Kang, Caixin, et al.
Published: (2023)