Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shabtay, Nimrod, Kimhi, Moshe, Spector, Artem, Haray, Sivan, Rivlin, Ehud, Baskin, Chaim, Giryes, Raja, Schwartz, Eli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CARES: Context-Aware Resolution Selector for VLMs
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
Deep Phase Coded Image Prior
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
PIP: Positional-encoding Image Prior
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2022)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2022)
Noisy Annotations in Semantic Segmentation
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
CLIMP: Contrastive Language-Image Mamba Pretraining
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2026)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2026)
Hysteresis Activation Function for Efficient Inference
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
Semi-Supervised Semantic Segmentation via Marginal Contextual Information
von: Kimhi, Moshe, et al.
Veröffentlicht: (2023)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2023)
$\mathbf{R}^3$: Reconstruction, Raw, and Rain: Deraining Directly in the Bayer Domain
von: Rothschild, Nate, et al.
Veröffentlicht: (2025)
von: Rothschild, Nate, et al.
Veröffentlicht: (2025)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
Robot Instance Segmentation with Few Annotations for Grasping
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
von: Blau, Tsachi, et al.
Veröffentlicht: (2024)
von: Blau, Tsachi, et al.
Veröffentlicht: (2024)
MAEDAY: MAE for few and zero shot AnomalY-Detection
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
von: Perek, Shaked, et al.
Veröffentlicht: (2026)
von: Perek, Shaked, et al.
Veröffentlicht: (2026)
Benchmarking Adversarial Patch Selection and Location
von: Kimhi, Shai, et al.
Veröffentlicht: (2025)
von: Kimhi, Shai, et al.
Veröffentlicht: (2025)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Leveraging Temporal Graph Networks Using Module Decoupling
von: Feldman, Or, et al.
Veröffentlicht: (2023)
von: Feldman, Or, et al.
Veröffentlicht: (2023)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
Looks Too Good To Be True: An Information-Theoretic Analysis of Hallucinations in Generative Restoration Models
von: Cohen, Regev, et al.
Veröffentlicht: (2024)
von: Cohen, Regev, et al.
Veröffentlicht: (2024)
Revisiting Node Affinity Prediction in Temporal Graphs
von: Feldman, Or, et al.
Veröffentlicht: (2025)
von: Feldman, Or, et al.
Veröffentlicht: (2025)
Leveraging Latents for Efficient Thermography Classification and Segmentation
von: Shor, Tamir, et al.
Veröffentlicht: (2024)
von: Shor, Tamir, et al.
Veröffentlicht: (2024)
Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
von: Taraday, Mitchell Keren, et al.
Veröffentlicht: (2025)
von: Taraday, Mitchell Keren, et al.
Veröffentlicht: (2025)
The CD-ROM Journal Literature: Where Do You Look?
von: Schwartz, Candy
Veröffentlicht: (1992)
von: Schwartz, Candy
Veröffentlicht: (1992)
Constitutional Governance in Metric Spaces
von: Shapiro, Ehud, et al.
Veröffentlicht: (2026)
von: Shapiro, Ehud, et al.
Veröffentlicht: (2026)
Grassroots Federation: Fair Democratic Governance at Scale
von: Shapiro, Ehud, et al.
Veröffentlicht: (2025)
von: Shapiro, Ehud, et al.
Veröffentlicht: (2025)
Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignment
von: Amoyal, Roy, et al.
Veröffentlicht: (2026)
von: Amoyal, Roy, et al.
Veröffentlicht: (2026)
Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges
von: Kosman, Eitan, et al.
Veröffentlicht: (2026)
von: Kosman, Eitan, et al.
Veröffentlicht: (2026)
On Adversarial Attacks In Acoustic Drone Localization
von: Shor, Tamir, et al.
Veröffentlicht: (2025)
von: Shor, Tamir, et al.
Veröffentlicht: (2025)
TEAM PILOT -- Learned Feasible Extendable Set of Dynamic MRI Acquisition Trajectories
von: Shor, Tamir, et al.
Veröffentlicht: (2024)
von: Shor, Tamir, et al.
Veröffentlicht: (2024)
Sparse patches adversarial attacks via extrapolating point-wise information
von: Nemcovsky, Yaniv, et al.
Veröffentlicht: (2024)
von: Nemcovsky, Yaniv, et al.
Veröffentlicht: (2024)
FLASH: Flexible Learning of Adaptive Sampling from History in Temporal Graph Neural Networks
von: Feldman, Or, et al.
Veröffentlicht: (2025)
von: Feldman, Or, et al.
Veröffentlicht: (2025)
Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
von: Galil, Ido, et al.
Veröffentlicht: (2025)
von: Galil, Ido, et al.
Veröffentlicht: (2025)
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
von: Serussi, Gabriele, et al.
Veröffentlicht: (2026)
von: Serussi, Gabriele, et al.
Veröffentlicht: (2026)
DifuzCam: Replacing Camera Lens with a Mask and a Diffusion Model
von: Yosef, Erez, et al.
Veröffentlicht: (2024)
von: Yosef, Erez, et al.
Veröffentlicht: (2024)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
von: Bar, Noga, et al.
Veröffentlicht: (2024)
von: Bar, Noga, et al.
Veröffentlicht: (2024)
Pruning at Initialization -- A Sketching Perspective
von: Bar, Noga, et al.
Veröffentlicht: (2023)
von: Bar, Noga, et al.
Veröffentlicht: (2023)
DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery
von: Khatib, Rajaei, et al.
Veröffentlicht: (2025)
von: Khatib, Rajaei, et al.
Veröffentlicht: (2025)
ZOQO: Zero-Order Quantized Optimization
von: Bar, Noga, et al.
Veröffentlicht: (2025)
von: Bar, Noga, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CARES: Context-Aware Resolution Selector for VLMs
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025) -
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025) -
Deep Phase Coded Image Prior
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024) -
PIP: Positional-encoding Image Prior
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2022) -
Noisy Annotations in Semantic Segmentation
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)