ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jewon, Shin, Wooksu, Yang, Seungmin, Song, Ki-Ung, Lim, DongUk, Kim, Jaeyeon, Kim, Tae-Ho, Kim, Bo-Kyeong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
Assessing the Answerability of Queries in Retrieval-Augmented Code Generation
by: Kim, Geonmin, et al.
Published: (2024)
by: Kim, Geonmin, et al.
Published: (2024)
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
by: Yeom, Seul-Ki, et al.
Published: (2024)
by: Yeom, Seul-Ki, et al.
Published: (2024)
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
by: Kim, Bo-Kyeong, et al.
Published: (2024)
by: Kim, Bo-Kyeong, et al.
Published: (2024)
Surgical Video Understanding with Label Interpolation
by: Kim, Garam, et al.
Published: (2025)
by: Kim, Garam, et al.
Published: (2025)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Environmental Understanding Vision-Language Model for Embodied Agent
by: Bang, Jinsik, et al.
Published: (2026)
by: Bang, Jinsik, et al.
Published: (2026)
Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth Prior
by: Hyoseok, Lee, et al.
Published: (2025)
by: Hyoseok, Lee, et al.
Published: (2025)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue Generation
by: Shin, Seungmin, et al.
Published: (2025)
by: Shin, Seungmin, et al.
Published: (2025)
Bi-MCQ: Reformulating Vision-Language Alignment for Negation Understanding
by: Kim, Tae Hun, et al.
Published: (2026)
by: Kim, Tae Hun, et al.
Published: (2026)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025)
by: Lee, YuEun, et al.
Published: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
by: Kim, Geewook, et al.
Published: (2024)
by: Kim, Geewook, et al.
Published: (2024)
Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History
by: Sung, Mingi, et al.
Published: (2024)
by: Sung, Mingi, et al.
Published: (2024)
Normalized Convolutional Neural Network
by: Kim, Dongsuk, et al.
Published: (2020)
by: Kim, Dongsuk, et al.
Published: (2020)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
by: Kim, Minkwan, et al.
Published: (2025)
by: Kim, Minkwan, et al.
Published: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
by: Lim, Su Hyeon, et al.
Published: (2024)
by: Lim, Su Hyeon, et al.
Published: (2024)
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
by: Simon, Marcel, et al.
Published: (2025)
by: Simon, Marcel, et al.
Published: (2025)
Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking
by: Puspitasari, Fachrina Dewi, et al.
Published: (2026)
by: Puspitasari, Fachrina Dewi, et al.
Published: (2026)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
by: Lim, Hyeonseok, et al.
Published: (2024)
by: Lim, Hyeonseok, et al.
Published: (2024)
Adaptive Teaching with Shared Classifier for Knowledge Distillation
by: Jang, Jaeyeon, et al.
Published: (2024)
by: Jang, Jaeyeon, et al.
Published: (2024)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
by: Yun, Heeseung, et al.
Published: (2025)
by: Yun, Heeseung, et al.
Published: (2025)
REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning
by: Lee, Seungmin, et al.
Published: (2026)
by: Lee, Seungmin, et al.
Published: (2026)
Lossless Token Merging Even Without Fine-Tuning in Vision Transformers
by: Lee, Jaeyeon, et al.
Published: (2025)
by: Lee, Jaeyeon, et al.
Published: (2025)
Factorized Multi-Resolution HashGrid for Efficient Neural Radiance Fields: Execution on Edge-Devices
by: Jun-Seong, Kim, et al.
Published: (2026)
by: Jun-Seong, Kim, et al.
Published: (2026)
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
by: Yeom, Seul-Ki, et al.
Published: (2026)
by: Yeom, Seul-Ki, et al.
Published: (2026)
GenQuery: Supporting Expressive Visual Search with Generative Models
by: Son, Kihoon, et al.
Published: (2023)
by: Son, Kihoon, et al.
Published: (2023)
IRASNet: Improved Feature-Level Clutter Reduction for Domain Generalized SAR-ATR
by: Jang, Oh-Tae, et al.
Published: (2024)
by: Jang, Oh-Tae, et al.
Published: (2024)
EdgeFusion: On-Device Text-to-Image Generation
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
QuadStretcher: A Forearm-Worn Skin Stretch Display for Bare-Hand Interaction in AR/VR
by: Kim, Taejun, et al.
Published: (2025)
by: Kim, Taejun, et al.
Published: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
HDR-NSFF: High Dynamic Range Neural Scene Flow Fields
by: Dong-Yeon, Shin, et al.
Published: (2026)
by: Dong-Yeon, Shin, et al.
Published: (2026)
Continuous Degradation Modeling via Latent Flow Matching for Real-World Super-Resolution
by: Kim, Hyeonjae, et al.
Published: (2026)
by: Kim, Hyeonjae, et al.
Published: (2026)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024)
by: Kim, Dongjin, et al.
Published: (2024)
PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation
by: Lim, Hyemin, et al.
Published: (2025)
by: Lim, Hyemin, et al.
Published: (2025)
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
by: Jung, Ji Hyeok, et al.
Published: (2024)
by: Jung, Ji Hyeok, et al.
Published: (2024)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
by: Shin, Dongjae, et al.
Published: (2024)
by: Shin, Dongjae, et al.
Published: (2024)
Similar Items
-
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025) -
Assessing the Answerability of Queries in Retrieval-Augmented Code Generation
by: Kim, Geonmin, et al.
Published: (2024) -
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
by: Yeom, Seul-Ki, et al.
Published: (2024) -
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
by: Kim, Bo-Kyeong, et al.
Published: (2024) -
Surgical Video Understanding with Label Interpolation
by: Kim, Garam, et al.
Published: (2025)