SinkTrack: Attention Sink based Context Anchoring for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xu, Chen, Guikun, Wang, Wenguan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scene Graph Generation with Role-Playing Large Language Models
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
von: Choi, Jiho, et al.
Veröffentlicht: (2026)
von: Choi, Jiho, et al.
Veröffentlicht: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
A Survey on 3D Gaussian Splatting
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
Neural Clustering based Visual Representation Learning
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
von: Chen, Minghan, et al.
Veröffentlicht: (2024)
von: Chen, Minghan, et al.
Veröffentlicht: (2024)
Attention Sinks in Diffusion Transformers: A Causal Analysis
von: Wu, Fangzheng, et al.
Veröffentlicht: (2026)
von: Wu, Fangzheng, et al.
Veröffentlicht: (2026)
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
von: Lu, Andrew, et al.
Veröffentlicht: (2025)
von: Lu, Andrew, et al.
Veröffentlicht: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
von: Wang, Yining, et al.
Veröffentlicht: (2025)
von: Wang, Yining, et al.
Veröffentlicht: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
von: Yoo, Suho, et al.
Veröffentlicht: (2026)
von: Yoo, Suho, et al.
Veröffentlicht: (2026)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
von: Ye, Bo, et al.
Veröffentlicht: (2026)
von: Ye, Bo, et al.
Veröffentlicht: (2026)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
Navigation Instruction Generation with BEV Perception and Large Language Models
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
Compositional Zero-shot Learning via Progressive Language-based Observations
von: Li, Lin, et al.
Veröffentlicht: (2023)
von: Li, Lin, et al.
Veröffentlicht: (2023)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
von: Li, Haodong, et al.
Veröffentlicht: (2026)
von: Li, Haodong, et al.
Veröffentlicht: (2026)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
von: Guo, Shanshan, et al.
Veröffentlicht: (2025)
von: Guo, Shanshan, et al.
Veröffentlicht: (2025)
Volumetric Environment Representation for Vision-Language Navigation
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Vision-Language Navigation with Energy-Based Policy
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
von: Li, Liulei, et al.
Veröffentlicht: (2024)
von: Li, Liulei, et al.
Veröffentlicht: (2024)
Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
von: Yi, Jung, et al.
Veröffentlicht: (2025)
von: Yi, Jung, et al.
Veröffentlicht: (2025)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
von: Liang, Chen, et al.
Veröffentlicht: (2022)
von: Liang, Chen, et al.
Veröffentlicht: (2022)
3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
SinkSAM-Net: Knowledge-Driven Self-Supervised Sinkhole Segmentation Using Topographic Priors and Segment Anything Model
von: Rafaeli, Osher, et al.
Veröffentlicht: (2024)
von: Rafaeli, Osher, et al.
Veröffentlicht: (2024)
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse Kernels
von: Feng, Tuo, et al.
Veröffentlicht: (2024)
von: Feng, Tuo, et al.
Veröffentlicht: (2024)
Learning Clustering-based Prototypes for Compositional Zero-shot Learning
von: Qu, Hongyu, et al.
Veröffentlicht: (2025)
von: Qu, Hongyu, et al.
Veröffentlicht: (2025)
A Survey of World Models for Autonomous Driving
von: Feng, Tuo, et al.
Veröffentlicht: (2025)
von: Feng, Tuo, et al.
Veröffentlicht: (2025)
Compositional Feature Augmentation for Unbiased Scene Graph Generation
von: Li, Lin, et al.
Veröffentlicht: (2023)
von: Li, Lin, et al.
Veröffentlicht: (2023)
Learning Human-Object Interaction as Groups
von: Hong, Jiajun, et al.
Veröffentlicht: (2025)
von: Hong, Jiajun, et al.
Veröffentlicht: (2025)
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
von: Chen, Mu, et al.
Veröffentlicht: (2025)
von: Chen, Mu, et al.
Veröffentlicht: (2025)
Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention
von: Zhang, Wenhu, et al.
Veröffentlicht: (2026)
von: Zhang, Wenhu, et al.
Veröffentlicht: (2026)
Decomposed Prototype Learning for Few-Shot Scene Graph Generation
von: Li, Xingchen, et al.
Veröffentlicht: (2023)
von: Li, Xingchen, et al.
Veröffentlicht: (2023)
Context-Aware Integration of Language and Visual References for Natural Language Tracking
von: Shao, Yanyan, et al.
Veröffentlicht: (2024)
von: Shao, Yanyan, et al.
Veröffentlicht: (2024)
Context-Aware Token Pruning and Discriminative Selective Attention for Transformer Tracking
von: Kugarajeevan, Janani, et al.
Veröffentlicht: (2025)
von: Kugarajeevan, Janani, et al.
Veröffentlicht: (2025)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
von: Hu, Ruina, et al.
Veröffentlicht: (2026)
von: Hu, Ruina, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scene Graph Generation with Role-Playing Large Language Models
von: Chen, Guikun, et al.
Veröffentlicht: (2024) -
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
von: Choi, Jiho, et al.
Veröffentlicht: (2026) -
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2025) -
A Survey on 3D Gaussian Splatting
von: Chen, Guikun, et al.
Veröffentlicht: (2024) -
Neural Clustering based Visual Representation Learning
von: Chen, Guikun, et al.
Veröffentlicht: (2024)