MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
Fuente:
arXiv
Saved in:
| Main Authors: | Gordon, Lucia, Belongie, Serge, Igel, Christian, Lang, Nico |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
by: Nedungadi, Vishal, et al.
Published: (2024)
by: Nedungadi, Vishal, et al.
Published: (2024)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023)
by: Enevoldsen, Philip, et al.
Published: (2023)
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Multimodal Fusion Strategies for Mapping Biophysical Landscape Features
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
POEM: Precise Object-level Editing via MLLM control
by: Schouten, Marco, et al.
Published: (2025)
by: Schouten, Marco, et al.
Published: (2025)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
by: Yu, Hong-Tao, et al.
Published: (2025)
by: Yu, Hong-Tao, et al.
Published: (2025)
Coarse-To-Fine Tensor Trains for Compact Visual Representations
by: Loeschcke, Sebastian, et al.
Published: (2024)
by: Loeschcke, Sebastian, et al.
Published: (2024)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026)
by: Schouten, Marco, et al.
Published: (2026)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation
by: Li, Jiao, et al.
Published: (2026)
by: Li, Jiao, et al.
Published: (2026)
SMART-PC: Skeletal Model Adaptation for Robust Test-Time Training in Point Clouds
by: Bahri, Ali, et al.
Published: (2025)
by: Bahri, Ali, et al.
Published: (2025)
Single Image Test-Time Adaptation via Multi-View Co-Training
by: Joshi, Smriti, et al.
Published: (2025)
by: Joshi, Smriti, et al.
Published: (2025)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
by: An, Zhaochong, et al.
Published: (2024)
by: An, Zhaochong, et al.
Published: (2024)
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
by: Deng, Qi, et al.
Published: (2024)
by: Deng, Qi, et al.
Published: (2024)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
Better Language Models Exhibit Higher Visual Alignment
by: Ruthardt, Jona, et al.
Published: (2024)
by: Ruthardt, Jona, et al.
Published: (2024)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
Space Rotation with Basis Transformation for Training-free Test-Time Adaptation
by: Ding, Chenhao, et al.
Published: (2025)
by: Ding, Chenhao, et al.
Published: (2025)
Test-Time Distillation for Continual Model Adaptation
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation
by: Li, Jiacheng, et al.
Published: (2025)
by: Li, Jiacheng, et al.
Published: (2025)
Realistic Test-Time Adaptation of Vision-Language Models
by: Zanella, Maxime, et al.
Published: (2025)
by: Zanella, Maxime, et al.
Published: (2025)
Bayesian Test-Time Adaptation for Vision-Language Models
by: Zhou, Lihua, et al.
Published: (2025)
by: Zhou, Lihua, et al.
Published: (2025)
Efficient Test-Time Adaptation of Vision-Language Models
by: Karmanov, Adilbek, et al.
Published: (2024)
by: Karmanov, Adilbek, et al.
Published: (2024)
Test-Time Model Adaptation for Quantized Neural Networks
by: Deng, Zeshuai, et al.
Published: (2025)
by: Deng, Zeshuai, et al.
Published: (2025)
Nip Rumors in the Bud: Retrieval-Guided Topic-Level Adaptation for Test-Time Fake News Video Detection
by: Lang, Jian, et al.
Published: (2026)
by: Lang, Jian, et al.
Published: (2026)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
by: Michael, Peter F., et al.
Published: (2025)
by: Michael, Peter F., et al.
Published: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation
by: Kim, Jihun, et al.
Published: (2026)
by: Kim, Jihun, et al.
Published: (2026)
Test-Time Adaptation of 3D Point Clouds via Denoising Diffusion Models
by: Dastmalchi, Hamidreza, et al.
Published: (2024)
by: Dastmalchi, Hamidreza, et al.
Published: (2024)
GS-Bias: Global-Spatial Bias Learner for Single-Image Test-Time Adaptation of Vision-Language Models
by: Huang, Zhaohong, et al.
Published: (2025)
by: Huang, Zhaohong, et al.
Published: (2025)
FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection
by: Zhao, Kaixiang, et al.
Published: (2026)
by: Zhao, Kaixiang, et al.
Published: (2026)
Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation
by: Noori, Mehrdad, et al.
Published: (2025)
by: Noori, Mehrdad, et al.
Published: (2025)
Assessing Neural Network Robustness via Adversarial Pivotal Tuning
by: Christensen, Peter Ebert, et al.
Published: (2022)
by: Christensen, Peter Ebert, et al.
Published: (2022)
CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation
by: Lafon, Marc, et al.
Published: (2025)
by: Lafon, Marc, et al.
Published: (2025)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
by: Zhu, Jinguo, et al.
Published: (2025)
by: Zhu, Jinguo, et al.
Published: (2025)
Similar Items
-
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
by: Nedungadi, Vishal, et al.
Published: (2024) -
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023) -
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
by: Jyhne, Sander Riisøen, et al.
Published: (2025) -
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024) -
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)