Saved in:
| Main Authors: | Choi, Joong Ho, Zhao, Jiayang, Appalla, Avani, Mukesh, Himansh, Vasani, Dhwanil, Qian, Boyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.02492 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
by: Choi, Joong Ho, et al.
Published: (2025)
by: Choi, Joong Ho, et al.
Published: (2025)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024)
by: Ju, Yeong-Joon, et al.
Published: (2024)
Naïve PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
by: Kim, Joong Ho, et al.
Published: (2026)
by: Kim, Joong Ho, et al.
Published: (2026)
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models
by: Liu, Yingen, et al.
Published: (2024)
by: Liu, Yingen, et al.
Published: (2024)
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
by: Zhao, Beidi, et al.
Published: (2025)
by: Zhao, Beidi, et al.
Published: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
Evolving Prompt Adaptation for Vision-Language Models
by: Zhang, Enming, et al.
Published: (2026)
by: Zhang, Enming, et al.
Published: (2026)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
by: Park, SoYoung, et al.
Published: (2025)
by: Park, SoYoung, et al.
Published: (2025)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
by: Yang, Zeyuan, et al.
Published: (2025)
by: Yang, Zeyuan, et al.
Published: (2025)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024)
by: Zhou, Qiji, et al.
Published: (2024)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
by: Chatterjee, Abhiroop, et al.
Published: (2025)
by: Chatterjee, Abhiroop, et al.
Published: (2025)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
by: Wang, Haoyu, et al.
Published: (2026)
by: Wang, Haoyu, et al.
Published: (2026)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
Language-Guided Invariance Probing of Vision-Language Models
by: Lee, Jae Joong
Published: (2025)
by: Lee, Jae Joong
Published: (2025)
3DTurboQuant: Training-Free Near-Optimal Quantization for 3D Reconstruction Models
by: Lee, Jae Joong
Published: (2026)
by: Lee, Jae Joong
Published: (2026)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
by: Li, Ruineng, et al.
Published: (2025)
by: Li, Ruineng, et al.
Published: (2025)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
MaeFuse: Transferring Omni Features with Pretrained Masked Autoencoders for Infrared and Visible Image Fusion via Guided Training
by: Li, Jiayang, et al.
Published: (2024)
by: Li, Jiayang, et al.
Published: (2024)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024)
by: Li, Yian, et al.
Published: (2024)
Chain of Time: In-Context Physical Simulation with Image Generation Models
by: Wang, YingQiao, et al.
Published: (2025)
by: Wang, YingQiao, et al.
Published: (2025)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
Semantically Aware UAV Landing Site Assessment from Remote Sensing Imagery via Multimodal Large Language Models
by: Hua, Chunliang, et al.
Published: (2026)
by: Hua, Chunliang, et al.
Published: (2026)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
by: Li, Jiaao, et al.
Published: (2025)
by: Li, Jiaao, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
by: Park, Jiwoong, et al.
Published: (2025)
by: Park, Jiwoong, et al.
Published: (2025)
From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
by: Du, Hang, et al.
Published: (2025)
by: Du, Hang, et al.
Published: (2025)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025)
by: Chen, Zisheng, et al.
Published: (2025)
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
by: Tu, Jiahang, et al.
Published: (2025)
by: Tu, Jiahang, et al.
Published: (2025)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
Similar Items
-
CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
by: Choi, Joong Ho, et al.
Published: (2025) -
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024) -
Naïve PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
by: Kim, Joong Ho, et al.
Published: (2026) -
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models
by: Liu, Yingen, et al.
Published: (2024) -
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
by: Xu, Zhe, et al.
Published: (2025)