CARES: Context-Aware Resolution Selector for VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kimhi, Moshe, Shabtay, Nimrod, Giryes, Raja, Baskin, Chaim, Schwartz, Eli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
PIP: Positional-encoding Image Prior
by: Shabtay, Nimrod, et al.
Published: (2022)
by: Shabtay, Nimrod, et al.
Published: (2022)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
Semi-Supervised Semantic Segmentation via Marginal Contextual Information
by: Kimhi, Moshe, et al.
Published: (2023)
by: Kimhi, Moshe, et al.
Published: (2023)
Deep Phase Coded Image Prior
by: Shabtay, Nimrod, et al.
Published: (2024)
by: Shabtay, Nimrod, et al.
Published: (2024)
CLIMP: Contrastive Language-Image Mamba Pretraining
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
Robot Instance Segmentation with Few Annotations for Grasping
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
Noisy Annotations in Semantic Segmentation
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
by: Galil, Ido, et al.
Published: (2025)
by: Galil, Ido, et al.
Published: (2025)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
by: Bar, Noga, et al.
Published: (2024)
by: Bar, Noga, et al.
Published: (2024)
$\mathbf{R}^3$: Reconstruction, Raw, and Rain: Deraining Directly in the Bayer Domain
by: Rothschild, Nate, et al.
Published: (2025)
by: Rothschild, Nate, et al.
Published: (2025)
Benchmarking Adversarial Patch Selection and Location
by: Kimhi, Shai, et al.
Published: (2025)
by: Kimhi, Shai, et al.
Published: (2025)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency
by: Dikter, Maor, et al.
Published: (2024)
by: Dikter, Maor, et al.
Published: (2024)
Sparse patches adversarial attacks via extrapolating point-wise information
by: Nemcovsky, Yaniv, et al.
Published: (2024)
by: Nemcovsky, Yaniv, et al.
Published: (2024)
Pruning at Initialization -- A Sketching Perspective
by: Bar, Noga, et al.
Published: (2023)
by: Bar, Noga, et al.
Published: (2023)
Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
by: Taraday, Mitchell Keren, et al.
Published: (2025)
by: Taraday, Mitchell Keren, et al.
Published: (2025)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
FedGSCA: Medical Federated Learning with Global Sample Selector and Client Adaptive Adjuster under Label Noise
by: Ye, Mengwen, et al.
Published: (2025)
by: Ye, Mengwen, et al.
Published: (2025)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
by: Augustin, Maximilian, et al.
Published: (2025)
by: Augustin, Maximilian, et al.
Published: (2025)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
The Power of Context: How Multimodality Improves Image Super-Resolution
by: Mei, Kangfu, et al.
Published: (2025)
by: Mei, Kangfu, et al.
Published: (2025)
Group Orthogonalization Regularization For Vision Models Adaptation and Robustness
by: Kurtz, Yoav, et al.
Published: (2023)
by: Kurtz, Yoav, et al.
Published: (2023)
ProtoSAM: One-Shot Medical Image Segmentation With Foundational Models
by: Ayzenberg, Lev, et al.
Published: (2024)
by: Ayzenberg, Lev, et al.
Published: (2024)
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
by: Ben-Ami, Dan, et al.
Published: (2026)
by: Ben-Ami, Dan, et al.
Published: (2026)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
by: Ding, Xi, et al.
Published: (2024)
by: Ding, Xi, et al.
Published: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024)
by: Micheli, Vincent, et al.
Published: (2024)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
DiffFinger: Advancing Synthetic Fingerprint Generation through Denoising Diffusion Probabilistic Models
by: Grabovski, Freddie, et al.
Published: (2024)
by: Grabovski, Freddie, et al.
Published: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
Similar Items
-
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026) -
PIP: Positional-encoding Image Prior
by: Shabtay, Nimrod, et al.
Published: (2022) -
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
by: Kimhi, Moshe, et al.
Published: (2025) -
Semi-Supervised Semantic Segmentation via Marginal Contextual Information
by: Kimhi, Moshe, et al.
Published: (2023) -
Deep Phase Coded Image Prior
by: Shabtay, Nimrod, et al.
Published: (2024)