When Do We Not Need Larger Vision Models?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Baifeng, Wu, Ziyang, Mao, Maolin, Wang, Xin, Darrell, Trevor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
Do We Need Reformer for Vision? An Experimental Comparison with Vision Transformers
von: Bellaj, Ali El, et al.
Veröffentlicht: (2025)
von: Bellaj, Ali El, et al.
Veröffentlicht: (2025)
Rethinking Patch Dependence for Masked Autoencoders
von: Fu, Letian, et al.
Veröffentlicht: (2024)
von: Fu, Letian, et al.
Veröffentlicht: (2024)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
MambaOut: Do We Really Need Mamba for Vision?
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
Humanoid Locomotion as Next Token Prediction
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?
von: Yan, Renye, et al.
Veröffentlicht: (2026)
von: Yan, Renye, et al.
Veröffentlicht: (2026)
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2025)
When Does Perceptual Alignment Benefit Vision Representations?
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2024)
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2024)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
von: Tao, Ming, et al.
Veröffentlicht: (2024)
von: Tao, Ming, et al.
Veröffentlicht: (2024)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Do Vision Language Models Need to Process Image Tokens?
von: Ghosh, Sambit, et al.
Veröffentlicht: (2026)
von: Ghosh, Sambit, et al.
Veröffentlicht: (2026)
Constantly Improving Image Models Need Constantly Improving Benchmarks
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
von: Tiong, Anthony Meng Huat, et al.
Veröffentlicht: (2024)
von: Tiong, Anthony Meng Huat, et al.
Veröffentlicht: (2024)
How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models
von: Hackel, Leonard, et al.
Veröffentlicht: (2026)
von: Hackel, Leonard, et al.
Veröffentlicht: (2026)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation
von: Kim, Taeyeong, et al.
Veröffentlicht: (2025)
von: Kim, Taeyeong, et al.
Veröffentlicht: (2025)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
REOrdering Patches Improves Vision Models
von: Kutscher, Declan, et al.
Veröffentlicht: (2025)
von: Kutscher, Declan, et al.
Veröffentlicht: (2025)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Hierarchical Invariance for Robust and Interpretable Vision Tasks at Larger Scales
von: Qi, Shuren, et al.
Veröffentlicht: (2024)
von: Qi, Shuren, et al.
Veröffentlicht: (2024)
Segment Anything without Supervision
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
Do We Really Need a Large Number of Visual Prompts?
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
von: Kuo, Shang-Jui Ray, et al.
Veröffentlicht: (2026)
von: Kuo, Shang-Jui Ray, et al.
Veröffentlicht: (2026)
Fast Image-based Neural Relighting with Translucency-Reflection Modeling
von: Zhu, Shizhan, et al.
Veröffentlicht: (2023)
von: Zhu, Shizhan, et al.
Veröffentlicht: (2023)
Vision Transformers Need More Than Registers
von: Shi, Cheng, et al.
Veröffentlicht: (2026)
von: Shi, Cheng, et al.
Veröffentlicht: (2026)
Learning to Grasp Anything by Playing with Random Toys
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
Readout Guidance: Learning Control from Diffusion Features
von: Luo, Grace, et al.
Veröffentlicht: (2023)
von: Luo, Grace, et al.
Veröffentlicht: (2023)
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
von: Lin, Liting, et al.
Veröffentlicht: (2024)
von: Lin, Liting, et al.
Veröffentlicht: (2024)
Larger than memory image processing
von: Sporring, Jon, et al.
Veröffentlicht: (2026)
von: Sporring, Jon, et al.
Veröffentlicht: (2026)
Finding Visual Task Vectors
von: Hojel, Alberto, et al.
Veröffentlicht: (2024)
von: Hojel, Alberto, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023) -
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023) -
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024) -
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024) -
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)