HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | SR, Nikitha, Mathur, Aradhya Neeraj, Menta, Tarun Ram, Jain, Rishabh, Sarkar, Mausoom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
by: SR, Nikitha, et al.
Published: (2024)
by: SR, Nikitha, et al.
Published: (2024)
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
by: Anand, Neeraj, et al.
Published: (2025)
by: Anand, Neeraj, et al.
Published: (2025)
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
by: Gopal, Varun, et al.
Published: (2026)
by: Gopal, Varun, et al.
Published: (2026)
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)
by: SR, Nikitha
Published: (2025)
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
by: Jadhav, Avadhoot, et al.
Published: (2025)
by: Jadhav, Avadhoot, et al.
Published: (2025)
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
by: Srivastava, Ashutosh, et al.
Published: (2024)
by: Srivastava, Ashutosh, et al.
Published: (2024)
MVGaussian: High-Fidelity text-to-3D Content Generation with Multi-View Guidance and Surface Densification
by: Pham, Phu, et al.
Published: (2024)
by: Pham, Phu, et al.
Published: (2024)
Curvy: A Parametric Cross-section based Surface Reconstruction
by: Mathur, Aradhya N., et al.
Published: (2024)
by: Mathur, Aradhya N., et al.
Published: (2024)
Evaluating Self-Correcting Vision Agents Through Quantitative and Qualitative Metrics
by: Dixit, Aradhya
Published: (2026)
by: Dixit, Aradhya
Published: (2026)
Multi-Level Feature Fusion Network for Lightweight Stereo Image Super-Resolution
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity
by: Inoue, Kotaro
Published: (2025)
by: Inoue, Kotaro
Published: (2025)
XFeat: Accelerated Features for Lightweight Image Matching
by: Potje, Guilherme, et al.
Published: (2024)
by: Potje, Guilherme, et al.
Published: (2024)
VarAD: Lightweight High-Resolution Image Anomaly Detection via Visual Autoregressive Modeling
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
Transformer-Progressive Mamba Network for Lightweight Image Super-Resolution
by: Guo, Sichen, et al.
Published: (2025)
by: Guo, Sichen, et al.
Published: (2025)
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025)
by: Liu, Wenyang, et al.
Published: (2025)
Streamlined Global and Local Features Combinator (SGLC) for High Resolution Image Dehazing
by: Benjdira, Bilel, et al.
Published: (2023)
by: Benjdira, Bilel, et al.
Published: (2023)
On the Effect of Image Resolution on Semantic Segmentation
by: Singh, Ritambhara, et al.
Published: (2024)
by: Singh, Ritambhara, et al.
Published: (2024)
Semantic Event Graphs for Long-Form Video Question Answering
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
Lightweight Adaptive Feature De-drifting for Compressed Image Classification
by: Peng, Long, et al.
Published: (2024)
by: Peng, Long, et al.
Published: (2024)
Semantic-Guided Global-Local Collaborative Networks for Lightweight Image Super-Resolution
by: Fan, Wanshu, et al.
Published: (2025)
by: Fan, Wanshu, et al.
Published: (2025)
Unifying Dimensions: A Linear Adaptive Approach to Lightweight Image Super-Resolution
by: Hu, Zhenyu, et al.
Published: (2024)
by: Hu, Zhenyu, et al.
Published: (2024)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation
by: Li, Qun, et al.
Published: (2022)
by: Li, Qun, et al.
Published: (2022)
Greit-HRNet: Grouped Lightweight High-Resolution Network for Human Pose Estimation
by: Han, Junjia
Published: (2024)
by: Han, Junjia
Published: (2024)
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-Resolution
by: Park, Karam, et al.
Published: (2025)
by: Park, Karam, et al.
Published: (2025)
CubeFormer: A Simple yet Effective Baseline for Lightweight Image Super-Resolution
by: Wang, Jikai, et al.
Published: (2024)
by: Wang, Jikai, et al.
Published: (2024)
Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
by: KV, Karthikeya
Published: (2025)
by: KV, Karthikeya
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition
by: Gupta, Rishabh, et al.
Published: (2026)
by: Gupta, Rishabh, et al.
Published: (2026)
Learning Accurate and Enriched Features for Stereo Image Super-Resolution
by: Gao, Hu, et al.
Published: (2024)
by: Gao, Hu, et al.
Published: (2024)
Feature Alignment with Equivariant Convolutions for Burst Image Super-Resolution
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
by: Khan, Shaharukh, et al.
Published: (2025)
by: Khan, Shaharukh, et al.
Published: (2025)
High Semantic Features for the Continual Learning of Complex Emotions: a Lightweight Solution
by: Geoffroy, Thibault, et al.
Published: (2025)
by: Geoffroy, Thibault, et al.
Published: (2025)
SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
by: Li, Linfei, et al.
Published: (2025)
by: Li, Linfei, et al.
Published: (2025)
LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection
by: Hao, Lei, et al.
Published: (2025)
by: Hao, Lei, et al.
Published: (2025)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025)
by: Hu, Jiewen, et al.
Published: (2025)
Similar Items
-
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
by: SR, Nikitha, et al.
Published: (2024) -
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
by: Anand, Neeraj, et al.
Published: (2025) -
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
by: Gopal, Varun, et al.
Published: (2026) -
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025) -
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)