$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
Fuente:
arXiv
Salvato in:
| Autori principali: | Maharana, Sarthak Kumar, Zhang, Baoming, Karlinsky, Leonid, Feris, Rogerio, Guo, Yunhui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
Latent Implicit Visual Reasoning
di: Li, Kelvin, et al.
Pubblicazione: (2025)
di: Li, Kelvin, et al.
Pubblicazione: (2025)
Not Just Change the Labels, Learn the Features: Watermarking Deep Neural Networks with Multi-View Data
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
STONE: A Submodular Optimization Framework for Active 3D Object Detection
di: Mao, Ruiyu, et al.
Pubblicazione: (2024)
di: Mao, Ruiyu, et al.
Pubblicazione: (2024)
Comparison Visual Instruction Tuning
di: Lin, Wei, et al.
Pubblicazione: (2024)
di: Lin, Wei, et al.
Pubblicazione: (2024)
Adaptive Memory Replay for Continual Learning
di: Smith, James Seale, et al.
Pubblicazione: (2024)
di: Smith, James Seale, et al.
Pubblicazione: (2024)
Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
di: Mao, Ruiyu, et al.
Pubblicazione: (2026)
di: Mao, Ruiyu, et al.
Pubblicazione: (2026)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
WATT: Weight Average Test-Time Adaptation of CLIP
di: Osowiechi, David, et al.
Pubblicazione: (2024)
di: Osowiechi, David, et al.
Pubblicazione: (2024)
Mitigating the ID-OOD Tradeoff in Open-Set Test-Time Adaptation
di: Zhao, Wenjie, et al.
Pubblicazione: (2026)
di: Zhao, Wenjie, et al.
Pubblicazione: (2026)
Sample- and Parameter-Efficient Auto-Regressive Image Models
di: Amrani, Elad, et al.
Pubblicazione: (2024)
di: Amrani, Elad, et al.
Pubblicazione: (2024)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
di: Yellinek, Nir, et al.
Pubblicazione: (2023)
di: Yellinek, Nir, et al.
Pubblicazione: (2023)
SafeFix: Targeted Model Repair via Controlled Image Generation
di: Xu, Ouyang, et al.
Pubblicazione: (2025)
di: Xu, Ouyang, et al.
Pubblicazione: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
Teaching VLMs to Localize Specific Objects from In-context Examples
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
CLIPArTT: Adaptation of CLIP to New Domains at Test Time
di: Hakim, Gustavo Adolfo Vargas, et al.
Pubblicazione: (2024)
di: Hakim, Gustavo Adolfo Vargas, et al.
Pubblicazione: (2024)
Activation Reward Models for Few-Shot Model Alignment
di: Chai, Tianning, et al.
Pubblicazione: (2025)
di: Chai, Tianning, et al.
Pubblicazione: (2025)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
di: Singh, Akshit, et al.
Pubblicazione: (2025)
di: Singh, Akshit, et al.
Pubblicazione: (2025)
Online Gaussian Test-Time Adaptation of Vision-Language Models
di: Fuchs, Clément, et al.
Pubblicazione: (2025)
di: Fuchs, Clément, et al.
Pubblicazione: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2025)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2025)
Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation
di: Chi, Zhixiang, et al.
Pubblicazione: (2025)
di: Chi, Zhixiang, et al.
Pubblicazione: (2025)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
di: Du, Yao, et al.
Pubblicazione: (2025)
di: Du, Yao, et al.
Pubblicazione: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
di: Araujo, Edson, et al.
Pubblicazione: (2025)
di: Araujo, Edson, et al.
Pubblicazione: (2025)
Source-Free Test-Time Adaptation For Online Surface-Defect Detection
di: Song, Yiran, et al.
Pubblicazione: (2024)
di: Song, Yiran, et al.
Pubblicazione: (2024)
Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
di: Dai, Qiyuan, et al.
Pubblicazione: (2025)
di: Dai, Qiyuan, et al.
Pubblicazione: (2025)
Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation
di: Gao, Junyu, et al.
Pubblicazione: (2023)
di: Gao, Junyu, et al.
Pubblicazione: (2023)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
di: Huang, Irene, et al.
Pubblicazione: (2024)
di: Huang, Irene, et al.
Pubblicazione: (2024)
Reshaping the Online Data Buffering and Organizing Mechanism for Continual Test-Time Adaptation
di: Zhu, Zhilin, et al.
Pubblicazione: (2024)
di: Zhu, Zhilin, et al.
Pubblicazione: (2024)
ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
di: Su, Linxiang, et al.
Pubblicazione: (2025)
di: Su, Linxiang, et al.
Pubblicazione: (2025)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
di: Huang, Brandon, et al.
Pubblicazione: (2025)
di: Huang, Brandon, et al.
Pubblicazione: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
di: Zeng, Ziyun, et al.
Pubblicazione: (2026)
di: Zeng, Ziyun, et al.
Pubblicazione: (2026)
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
di: Liu, Jiaming, et al.
Pubblicazione: (2023)
di: Liu, Jiaming, et al.
Pubblicazione: (2023)
COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation
di: Monga, Munish, et al.
Pubblicazione: (2024)
di: Monga, Munish, et al.
Pubblicazione: (2024)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
di: Alyami, Sarah, et al.
Pubblicazione: (2025)
di: Alyami, Sarah, et al.
Pubblicazione: (2025)
Domain-Specific Block Selection and Paired-View Pseudo-Labeling for Online Test-Time Adaptation
di: Yu, Yeonguk, et al.
Pubblicazione: (2024)
di: Yu, Yeonguk, et al.
Pubblicazione: (2024)
Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation
di: Wu, Xiangyu, et al.
Pubblicazione: (2026)
di: Wu, Xiangyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024) -
Latent Implicit Visual Reasoning
di: Li, Kelvin, et al.
Pubblicazione: (2025) -
Not Just Change the Labels, Learn the Features: Watermarking Deep Neural Networks with Multi-View Data
di: Li, Yuxuan, et al.
Pubblicazione: (2024) -
STONE: A Submodular Optimization Framework for Active 3D Object Detection
di: Mao, Ruiyu, et al.
Pubblicazione: (2024) -
Comparison Visual Instruction Tuning
di: Lin, Wei, et al.
Pubblicazione: (2024)