Efficient Whole Slide Pathology VQA via Token Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Weimin, Hu, Qingqiao, Qi, Kehan, Shi, Zhan, Huang, Wentao, Gupta, Saumya, Chen, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
by: Hu, Qingqiao, et al.
Published: (2025)
by: Hu, Qingqiao, et al.
Published: (2025)
Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
by: Qi, Kehan, et al.
Published: (2025)
by: Qi, Kehan, et al.
Published: (2025)
Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning
by: Huang, Wentao, et al.
Published: (2026)
by: Huang, Wentao, et al.
Published: (2026)
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
by: Xu, Meilong, et al.
Published: (2026)
by: Xu, Meilong, et al.
Published: (2026)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
by: Chen, Pingyi, et al.
Published: (2023)
by: Chen, Pingyi, et al.
Published: (2023)
Hard Negative Sample Mining for Whole Slide Image Classification
by: Huang, Wentao, et al.
Published: (2024)
by: Huang, Wentao, et al.
Published: (2024)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
by: Liang, Yuci, et al.
Published: (2024)
by: Liang, Yuci, et al.
Published: (2024)
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
by: Pan, Yiming, et al.
Published: (2026)
by: Pan, Yiming, et al.
Published: (2026)
PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide Image VQA
by: Yang, Chunze, et al.
Published: (2026)
by: Yang, Chunze, et al.
Published: (2026)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
by: Li, Yangfu, et al.
Published: (2025)
by: Li, Yangfu, et al.
Published: (2025)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
by: Liu, Xinxin, et al.
Published: (2026)
by: Liu, Xinxin, et al.
Published: (2026)
Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis
by: Yang, Zhidong, et al.
Published: (2025)
by: Yang, Zhidong, et al.
Published: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
by: Kim, Sangwook, et al.
Published: (2025)
by: Kim, Sangwook, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
by: Li, Xue, et al.
Published: (2025)
by: Li, Xue, et al.
Published: (2025)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025)
by: Ge, Jinchao, et al.
Published: (2025)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
Diagnostic Text-guided Representation Learning in Hierarchical Classification for Pathological Whole Slide Image
by: Li, Jiawen, et al.
Published: (2024)
by: Li, Jiawen, et al.
Published: (2024)
SEW: Self-calibration Enhanced Whole Slide Pathology Image Analysis
by: Luo, Haoming, et al.
Published: (2024)
by: Luo, Haoming, et al.
Published: (2024)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
by: Chen, Ying, et al.
Published: (2024)
by: Chen, Ying, et al.
Published: (2024)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
by: Mo, Wentao, et al.
Published: (2024)
by: Mo, Wentao, et al.
Published: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Backdooring Vision-Language Models with Out-Of-Distribution Data
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
Efficient Quality Control of Whole Slide Pathology Images with Human-in-the-loop Training
by: Patil, Abhijeet, et al.
Published: (2024)
by: Patil, Abhijeet, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics
by: Zhang, Yupei, et al.
Published: (2025)
by: Zhang, Yupei, et al.
Published: (2025)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
by: Jiang, Titong, et al.
Published: (2025)
by: Jiang, Titong, et al.
Published: (2025)
Similar Items
-
LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
by: Hu, Qingqiao, et al.
Published: (2025) -
Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
by: Qi, Kehan, et al.
Published: (2025) -
Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning
by: Huang, Wentao, et al.
Published: (2026) -
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
by: Xu, Meilong, et al.
Published: (2026) -
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)