VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Hengbo, Jin, Shengjie, Ma, Yanbiao, Lu, Zhiwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparsity- and Hybridity-Inspired Visual Parameter-Efficient Fine-Tuning for Medical Diagnosis
von: Liu, Mingyuan, et al.
Veröffentlicht: (2024)
von: Liu, Mingyuan, et al.
Veröffentlicht: (2024)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
Geometric Origins of Bias in Deep Neural Networks: A Human Visual System Perspective
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
Learnable Sparsity for Vision Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2024)
von: Zhang, Yang, et al.
Veröffentlicht: (2024)
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Sparsity Meets Similarity: Leveraging Long-Tail Distribution for Dynamic Optimized Token Representation in Multimodal Large Language Models
von: Yu, Gaotong, et al.
Veröffentlicht: (2024)
von: Yu, Gaotong, et al.
Veröffentlicht: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks
von: Jin, Jing, et al.
Veröffentlicht: (2026)
von: Jin, Jing, et al.
Veröffentlicht: (2026)
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
von: Zuo, Jing, et al.
Veröffentlicht: (2026)
von: Zuo, Jing, et al.
Veröffentlicht: (2026)
Efficient Modulation for Vision Networks
von: Ma, Xu, et al.
Veröffentlicht: (2024)
von: Ma, Xu, et al.
Veröffentlicht: (2024)
Compositional Attribute Imbalance in Vision Datasets
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
von: Wang, Yu, et al.
Veröffentlicht: (2026)
von: Wang, Yu, et al.
Veröffentlicht: (2026)
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
von: Chen, Zeyuan, et al.
Veröffentlicht: (2025)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2025)
Question Aware Vision Transformer for Multimodal Reasoning
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
von: Ai, Hao, et al.
Veröffentlicht: (2025)
von: Ai, Hao, et al.
Veröffentlicht: (2025)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
BabyVision: Visual Reasoning Beyond Language
von: Chen, Liang, et al.
Veröffentlicht: (2026)
von: Chen, Liang, et al.
Veröffentlicht: (2026)
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
von: Majeedi, Abrar, et al.
Veröffentlicht: (2026)
von: Majeedi, Abrar, et al.
Veröffentlicht: (2026)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
Exploring Beyond Logits: Hierarchical Dynamic Labeling Based on Embeddings for Semi-Supervised Classification
von: Ma, Yanbiao, et al.
Veröffentlicht: (2024)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2024)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2023)
von: Jin, Yang, et al.
Veröffentlicht: (2023)
Towards General Multimodal Visual Tracking
von: Lu, Andong, et al.
Veröffentlicht: (2025)
von: Lu, Andong, et al.
Veröffentlicht: (2025)
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
Pursuing Better Decision Boundaries for Long-Tailed Object Detection via Category Information Amount
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
Deep Feature Gaussian Processes for Single-Scene Aerosol Optical Depth Reconstruction
von: Liu, Shengjie, et al.
Veröffentlicht: (2024)
von: Liu, Shengjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sparsity- and Hybridity-Inspired Visual Parameter-Efficient Fine-Tuning for Medical Diagnosis
von: Liu, Mingyuan, et al.
Veröffentlicht: (2024) -
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
von: Khaki, Samir, et al.
Veröffentlicht: (2025) -
Geometric Origins of Bias in Deep Neural Networks: A Human Visual System Perspective
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025) -
Learnable Sparsity for Vision Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2024) -
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)