A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Jaeseong, Choi, Yeeun, Choi, Heechan, Kim, Hanjung, Kim, Seonjoo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fourier-Guided Attention Upsampling for Image Super-Resolution
por: Choi, Daejune, et al.
Publicado: (2025)
por: Choi, Daejune, et al.
Publicado: (2025)
MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
por: Choi, Changho, et al.
Publicado: (2025)
por: Choi, Changho, et al.
Publicado: (2025)
Training-Free Restoration of Pruned Neural Networks
por: Lee, Keonho, et al.
Publicado: (2025)
por: Lee, Keonho, et al.
Publicado: (2025)
Addressing Diverging Training Costs using BEVRestore for High-resolution Bird's Eye View Map Construction
por: Kim, Minsu, et al.
Publicado: (2024)
por: Kim, Minsu, et al.
Publicado: (2024)
Colorful Cutout: Enhancing Image Data Augmentation with Curriculum Learning
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
por: Lee, Jaeseong, et al.
Publicado: (2024)
por: Lee, Jaeseong, et al.
Publicado: (2024)
LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights
por: Castells, Thibault, et al.
Publicado: (2024)
por: Castells, Thibault, et al.
Publicado: (2024)
FreeVA: Offline MLLM as Training-Free Video Assistant
por: Wu, Wenhao
Publicado: (2024)
por: Wu, Wenhao
Publicado: (2024)
Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention
por: Park, Jeonghoon, et al.
Publicado: (2025)
por: Park, Jeonghoon, et al.
Publicado: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
por: Lee, Hankyeol, et al.
Publicado: (2025)
por: Lee, Hankyeol, et al.
Publicado: (2025)
EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams
por: Kim, Jaeseong, et al.
Publicado: (2026)
por: Kim, Jaeseong, et al.
Publicado: (2026)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
por: Chi, Donghwan, et al.
Publicado: (2025)
por: Chi, Donghwan, et al.
Publicado: (2025)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
por: Kim, Sanghyun, et al.
Publicado: (2024)
por: Kim, Sanghyun, et al.
Publicado: (2024)
Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion
por: Kim, Sanghyun, et al.
Publicado: (2024)
por: Kim, Sanghyun, et al.
Publicado: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
por: Shin, Hyunjune, et al.
Publicado: (2024)
por: Shin, Hyunjune, et al.
Publicado: (2024)
Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentation
por: Oh, Gyeongrok, et al.
Publicado: (2025)
por: Oh, Gyeongrok, et al.
Publicado: (2025)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
por: Lee, Eunsoo, et al.
Publicado: (2026)
por: Lee, Eunsoo, et al.
Publicado: (2026)
DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models
por: Chae, Daewon, et al.
Publicado: (2025)
por: Chae, Daewon, et al.
Publicado: (2025)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
por: Han, Sangmin, et al.
Publicado: (2025)
por: Han, Sangmin, et al.
Publicado: (2025)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
por: Kim, Jihwan, et al.
Publicado: (2024)
por: Kim, Jihwan, et al.
Publicado: (2024)
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
por: Oh, Gyeongrok, et al.
Publicado: (2025)
por: Oh, Gyeongrok, et al.
Publicado: (2025)
FreeSliders: Training-Free, Modality-Agnostic Concept Sliders for Fine-Grained Diffusion Control in Images, Audio, and Video
por: Ezra, Rotem, et al.
Publicado: (2025)
por: Ezra, Rotem, et al.
Publicado: (2025)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
por: Kim, Jihyeon, et al.
Publicado: (2026)
por: Kim, Jihyeon, et al.
Publicado: (2026)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
por: Lee, Sangin, et al.
Publicado: (2026)
por: Lee, Sangin, et al.
Publicado: (2026)
RoAD Benchmark: How LiDAR Models Fail under Coupled Domain Shifts and Label Evolution
por: Lee, Subeen, et al.
Publicado: (2026)
por: Lee, Subeen, et al.
Publicado: (2026)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2024)
por: Woo, Sangmin, et al.
Publicado: (2024)
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
por: Kim, Kyeong Seon, et al.
Publicado: (2026)
por: Kim, Kyeong Seon, et al.
Publicado: (2026)
Text-Guided Variational Image Generation for Industrial Anomaly Detection and Segmentation
por: Lee, Mingyu, et al.
Publicado: (2024)
por: Lee, Mingyu, et al.
Publicado: (2024)
Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
por: Chang, Jinho, et al.
Publicado: (2025)
por: Chang, Jinho, et al.
Publicado: (2025)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
por: Choi, Jae Young, et al.
Publicado: (2026)
por: Choi, Jae Young, et al.
Publicado: (2026)
Reciprocal Attention Mixing Transformer for Lightweight Image Restoration
por: Choi, Haram, et al.
Publicado: (2023)
por: Choi, Haram, et al.
Publicado: (2023)
CIPHER: Counterfeit Image Pattern High-level Examination via Representation
por: Kim, Kyeonghun, et al.
Publicado: (2026)
por: Kim, Kyeonghun, et al.
Publicado: (2026)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
por: Kim, Jaemin, et al.
Publicado: (2024)
por: Kim, Jaemin, et al.
Publicado: (2024)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
por: Jeong, Suchae, et al.
Publicado: (2025)
por: Jeong, Suchae, et al.
Publicado: (2025)
Training-Free Label Space Alignment for Universal Domain Adaptation
por: Lee, Dujin, et al.
Publicado: (2025)
por: Lee, Dujin, et al.
Publicado: (2025)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
por: Min, Jeongho, et al.
Publicado: (2025)
por: Min, Jeongho, et al.
Publicado: (2025)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
por: Park, Jiho, et al.
Publicado: (2025)
por: Park, Jiho, et al.
Publicado: (2025)
What Happens When: Learning Temporal Orders of Events in Videos
por: Ahn, Daechul, et al.
Publicado: (2025)
por: Ahn, Daechul, et al.
Publicado: (2025)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
por: Kwon, Mincheol, et al.
Publicado: (2026)
por: Kwon, Mincheol, et al.
Publicado: (2026)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
por: Cho, Jungbin, et al.
Publicado: (2025)
por: Cho, Jungbin, et al.
Publicado: (2025)
Ejemplares similares
-
Fourier-Guided Attention Upsampling for Image Super-Resolution
por: Choi, Daejune, et al.
Publicado: (2025) -
MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
por: Choi, Changho, et al.
Publicado: (2025) -
Training-Free Restoration of Pruned Neural Networks
por: Lee, Keonho, et al.
Publicado: (2025) -
Addressing Diverging Training Costs using BEVRestore for High-resolution Bird's Eye View Map Construction
por: Kim, Minsu, et al.
Publicado: (2024) -
Colorful Cutout: Enhancing Image Data Augmentation with Curriculum Learning
por: Choi, Juhwan, et al.
Publicado: (2024)