Prism: Spectral-Aware Block-Sparse Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xinghao, Wang, Pengyu, Liu, Xiaoran, Liu, Fangxu, Chu, Jason, Song, Kai, Qiu, Xipeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sparser Block-Sparse Attention via Token Permutation
por: Wang, Xinghao, et al.
Publicado: (2025)
por: Wang, Xinghao, et al.
Publicado: (2025)
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
por: Wang, Xinghao, et al.
Publicado: (2024)
por: Wang, Xinghao, et al.
Publicado: (2024)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
por: Zong, Yi, et al.
Publicado: (2024)
por: Zong, Yi, et al.
Publicado: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
por: Chen, Zhuokun, et al.
Publicado: (2026)
por: Chen, Zhuokun, et al.
Publicado: (2026)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
por: Wang, Xinghao, et al.
Publicado: (2025)
por: Wang, Xinghao, et al.
Publicado: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
por: Zhang, Haowei, et al.
Publicado: (2026)
por: Zhang, Haowei, et al.
Publicado: (2026)
LatentLLM: Attention-Aware Joint Tensor Compression
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
por: Liu, Peiju, et al.
Publicado: (2026)
por: Liu, Peiju, et al.
Publicado: (2026)
XAttention: Block Sparse Attention with Antidiagonal Scoring
por: Xu, Ruyi, et al.
Publicado: (2025)
por: Xu, Ruyi, et al.
Publicado: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
por: Jiang, Chen, et al.
Publicado: (2023)
por: Jiang, Chen, et al.
Publicado: (2023)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
por: Fan, Xiaoran, et al.
Publicado: (2026)
por: Fan, Xiaoran, et al.
Publicado: (2026)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
por: Zhu, Chenming, et al.
Publicado: (2024)
por: Zhu, Chenming, et al.
Publicado: (2024)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
por: Chen, Wenting, et al.
Publicado: (2024)
por: Chen, Wenting, et al.
Publicado: (2024)
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
por: Zhang, Yuyi, et al.
Publicado: (2025)
por: Zhang, Yuyi, et al.
Publicado: (2025)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
por: Gao, Xiangbo, et al.
Publicado: (2026)
por: Gao, Xiangbo, et al.
Publicado: (2026)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
por: Xie, Xudong, et al.
Publicado: (2024)
por: Xie, Xudong, et al.
Publicado: (2024)
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
por: Wu, Jiarui, et al.
Publicado: (2025)
por: Wu, Jiarui, et al.
Publicado: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
por: Zhang, Shiduo, et al.
Publicado: (2024)
por: Zhang, Shiduo, et al.
Publicado: (2024)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
por: Chen, Yi-Chia, et al.
Publicado: (2024)
por: Chen, Yi-Chia, et al.
Publicado: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
por: Li, Yanwei, et al.
Publicado: (2024)
por: Li, Yanwei, et al.
Publicado: (2024)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
por: Lin, Rui, et al.
Publicado: (2026)
por: Lin, Rui, et al.
Publicado: (2026)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
por: Chen, Xinlong, et al.
Publicado: (2025)
por: Chen, Xinlong, et al.
Publicado: (2025)
S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing
por: Wang, Guangzhi, et al.
Publicado: (2024)
por: Wang, Guangzhi, et al.
Publicado: (2024)
Evolutionary Negative Module Pruning for Better LoRA Merging
por: Cao, Anda, et al.
Publicado: (2026)
por: Cao, Anda, et al.
Publicado: (2026)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
por: Feng, Sicheng, et al.
Publicado: (2025)
por: Feng, Sicheng, et al.
Publicado: (2025)
Fostering Video Reasoning via Next-Event Prediction
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
por: An, Wenbin, et al.
Publicado: (2024)
por: An, Wenbin, et al.
Publicado: (2024)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
por: Shen, Li-Cheng, et al.
Publicado: (2025)
por: Shen, Li-Cheng, et al.
Publicado: (2025)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
por: Liu, Yuhan, et al.
Publicado: (2025)
por: Liu, Yuhan, et al.
Publicado: (2025)
Interpretable and Sparse Linear Attention with Decoupled Membership-Subspace Modeling via MCR2 Objective
por: Liu, Tianyuan, et al.
Publicado: (2026)
por: Liu, Tianyuan, et al.
Publicado: (2026)
VideoPrism: A Foundational Visual Encoder for Video Understanding
por: Zhao, Long, et al.
Publicado: (2024)
por: Zhao, Long, et al.
Publicado: (2024)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
por: Huang, Yidong, et al.
Publicado: (2025)
por: Huang, Yidong, et al.
Publicado: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention
por: Xu, Fengyi, et al.
Publicado: (2025)
por: Xu, Fengyi, et al.
Publicado: (2025)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
por: Chandu, Khyathi Raghavi, et al.
Publicado: (2024)
por: Chandu, Khyathi Raghavi, et al.
Publicado: (2024)
On Data Synthesis and Post-training for Visual Abstract Reasoning
por: Zhu, Ke, et al.
Publicado: (2025)
por: Zhu, Ke, et al.
Publicado: (2025)
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
por: Cui, Guofeng, et al.
Publicado: (2025)
por: Cui, Guofeng, et al.
Publicado: (2025)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
por: Chen, Xinyu, et al.
Publicado: (2026)
por: Chen, Xinyu, et al.
Publicado: (2026)
Ejemplares similares
-
Sparser Block-Sparse Attention via Token Permutation
por: Wang, Xinghao, et al.
Publicado: (2025) -
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
por: Wang, Xinghao, et al.
Publicado: (2024) -
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
por: Zong, Yi, et al.
Publicado: (2024) -
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
por: Chen, Zhuokun, et al.
Publicado: (2026) -
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
por: Wang, Xinghao, et al.
Publicado: (2025)