Revisiting Adversarial Training at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zeyu, Li, Xianhang, Zhu, Hongru, Xie, Cihang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
Rejuvenating image-GPT as Strong Visual Representation Learners
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
On the Adversarial Robustness of Camera-based 3D Object Detection
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2023)
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2023)
Scaling White-Box Transformers for Vision
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
von: Li, Zichao, et al.
Veröffentlicht: (2024)
von: Li, Zichao, et al.
Veröffentlicht: (2024)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
L2B: Learning to Bootstrap Robust Models for Combating Label Noise
von: Zhou, Yuyin, et al.
Veröffentlicht: (2022)
von: Zhou, Yuyin, et al.
Veröffentlicht: (2022)
Revisiting Adversarial Training under Long-Tailed Distributions
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
Revisiting Adversarial Training under Hyperspectral Image
von: Zhang, Weihua, et al.
Veröffentlicht: (2025)
von: Zhang, Weihua, et al.
Veröffentlicht: (2025)
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
Autoregressive Pretraining with Mamba in Vision
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Steering Video Diffusion Transformers with Massive Activations
von: Cheng, Xianhang, et al.
Veröffentlicht: (2026)
von: Cheng, Xianhang, et al.
Veröffentlicht: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
Revisiting Min-Max Optimization Problem in Adversarial Training
von: Ahmadi, Sina Hajer, et al.
Veröffentlicht: (2024)
von: Ahmadi, Sina Hajer, et al.
Veröffentlicht: (2024)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
MapNeXt: Revisiting Training and Scaling Practices for Online Vectorized HD Map Construction
von: Li, Toyota
Veröffentlicht: (2024)
von: Li, Toyota
Veröffentlicht: (2024)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training
von: Wang, Yanyun, et al.
Veröffentlicht: (2026)
von: Wang, Yanyun, et al.
Veröffentlicht: (2026)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better Generalization
von: Hou, Saihui, et al.
Veröffentlicht: (2024)
von: Hou, Saihui, et al.
Veröffentlicht: (2024)
Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
von: Zheng, Junhao, et al.
Veröffentlicht: (2025)
von: Zheng, Junhao, et al.
Veröffentlicht: (2025)
ARFlow: Autoregressive Flow with Hybrid Linear Attention
von: Hui, Mude, et al.
Veröffentlicht: (2025)
von: Hui, Mude, et al.
Veröffentlicht: (2025)
SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models
von: Chen, Hesen, et al.
Veröffentlicht: (2025)
von: Chen, Hesen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025) -
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024) -
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025) -
Rejuvenating image-GPT as Strong Visual Representation Learners
von: Ren, Sucheng, et al.
Veröffentlicht: (2023) -
On the Adversarial Robustness of Camera-based 3D Object Detection
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2023)