Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Yun, Jungmin, Kim, Mihyeon, Kim, Youngbin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
by: Kim, Mihyeon, et al.
Published: (2025)
by: Kim, Mihyeon, et al.
Published: (2025)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
by: Jin, Kyohoon, et al.
Published: (2025)
by: Jin, Kyohoon, et al.
Published: (2025)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
by: Son, Jaemin, et al.
Published: (2025)
by: Son, Jaemin, et al.
Published: (2025)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
From Ground Trust to Truth: Disparities in Offensive Language Judgments on Contemporary Korean Political Discourse
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference
by: Liao, Ruqi, et al.
Published: (2024)
by: Liao, Ruqi, et al.
Published: (2024)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
by: Zhang, Yong, et al.
Published: (2025)
by: Zhang, Yong, et al.
Published: (2025)
Saliency-driven Dynamic Token Pruning for Large Language Models
by: Tao, Yao, et al.
Published: (2025)
by: Tao, Yao, et al.
Published: (2025)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
by: Figliolia, Tomas, et al.
Published: (2025)
by: Figliolia, Tomas, et al.
Published: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
by: Jin, Kyohoon, et al.
Published: (2024)
by: Jin, Kyohoon, et al.
Published: (2024)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
by: Kim, San, et al.
Published: (2025)
by: Kim, San, et al.
Published: (2025)
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
by: Liu, Deyuan, et al.
Published: (2024)
by: Liu, Deyuan, et al.
Published: (2024)
Lossless Token Sequence Compression via Meta-Tokens
by: Harvill, John, et al.
Published: (2025)
by: Harvill, John, et al.
Published: (2025)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRA
by: Jung, Chanjoo, et al.
Published: (2025)
by: Jung, Chanjoo, et al.
Published: (2025)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
by: Choi, Minsik, et al.
Published: (2025)
by: Choi, Minsik, et al.
Published: (2025)
Control Token with Dense Passage Retrieval
by: Lee, Juhwan, et al.
Published: (2024)
by: Lee, Juhwan, et al.
Published: (2024)
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
by: Fang, Yixiong, et al.
Published: (2025)
by: Fang, Yixiong, et al.
Published: (2025)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
by: Lee, Nakyung, et al.
Published: (2025)
by: Lee, Nakyung, et al.
Published: (2025)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
by: Zhou, Longsheng, et al.
Published: (2026)
by: Zhou, Longsheng, et al.
Published: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
by: Chen, Lizhe, et al.
Published: (2025)
by: Chen, Lizhe, et al.
Published: (2025)
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026)
by: Kim, Ireh, et al.
Published: (2026)
Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
by: Kim, Singon
Published: (2025)
by: Kim, Singon
Published: (2025)
Incorporating Domain Knowledge into Materials Tokenization
by: Oh, Yerim, et al.
Published: (2025)
by: Oh, Yerim, et al.
Published: (2025)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
by: Kim, Jongwook, et al.
Published: (2025)
by: Kim, Jongwook, et al.
Published: (2025)
Similar Items
-
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
by: Kim, Mihyeon, et al.
Published: (2025) -
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
by: Jin, Kyohoon, et al.
Published: (2025) -
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025) -
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
by: Choi, Juhwan, et al.
Published: (2024) -
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
by: Son, Jaemin, et al.
Published: (2025)