Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhisong, Wang, Yan, Huang, Xinting, Fang, Tianqing, Zhang, Hongming, Deng, Chenlong, Li, Shuaiyi, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2024)
von: Deng, Chenlong, et al.
Veröffentlicht: (2024)
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2025)
von: Deng, Chenlong, et al.
Veröffentlicht: (2025)
InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
von: Li, Shuaiyi, et al.
Veröffentlicht: (2025)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2025)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
von: Li, Shuaiyi, et al.
Veröffentlicht: (2026)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2026)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
von: Cai, Deng, et al.
Veröffentlicht: (2024)
von: Cai, Deng, et al.
Veröffentlicht: (2024)
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
von: Hu, Minda, et al.
Veröffentlicht: (2025)
von: Hu, Minda, et al.
Veröffentlicht: (2025)
Pre-training Everywhere: Parameter-Efficient Fine-Tuning for Medical Image Analysis via Target Parameter Pre-training
von: Lei, Xingliang, et al.
Veröffentlicht: (2024)
von: Lei, Xingliang, et al.
Veröffentlicht: (2024)
Long-Context Language Modeling with Parallel Context Encoding
von: Yen, Howard, et al.
Veröffentlicht: (2024)
von: Yen, Howard, et al.
Veröffentlicht: (2024)
On-the-fly Denoising for Data Augmentation in Natural Language Understanding
von: Fang, Tianqing, et al.
Veröffentlicht: (2022)
von: Fang, Tianqing, et al.
Veröffentlicht: (2022)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Atomic Calibration of LLMs in Long-Form Generations
von: Zhang, Caiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Caiqi, et al.
Veröffentlicht: (2024)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
LoGU: Long-form Generation with Uncertainty Expressions
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Do Pre-trained Vision-Language Models Encode Object States?
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
Multiple-Debias: A Full-process Debiasing Method for Multilingual Pre-trained Language Models
von: Liang, Haoyu, et al.
Veröffentlicht: (2026)
von: Liang, Haoyu, et al.
Veröffentlicht: (2026)
Adaptive Kernel Density Estimation with Pre-training
von: Zhang, Ruitong, et al.
Veröffentlicht: (2026)
von: Zhang, Ruitong, et al.
Veröffentlicht: (2026)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
SEER: Spectral Entropy Encoding of Roles for Context-Aware Attention-Based Design Pattern Detection
von: Houichime, Tarik, et al.
Veröffentlicht: (2026)
von: Houichime, Tarik, et al.
Veröffentlicht: (2026)
PaPaformer: Language Model from Pre-trained Parallel Paths
von: Tapaninaho, Joonas, et al.
Veröffentlicht: (2025)
von: Tapaninaho, Joonas, et al.
Veröffentlicht: (2025)
Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA
von: Qin, Laiqiao, et al.
Veröffentlicht: (2024)
von: Qin, Laiqiao, et al.
Veröffentlicht: (2024)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
von: Wang, Zhaowei, et al.
Veröffentlicht: (2023)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2023)
YOLOv8‐based framework for medial temporal lobe atrophy grading
von: Tianqing Deng, et al.
Veröffentlicht: (2025)
von: Tianqing Deng, et al.
Veröffentlicht: (2025)
Entropy-Based Data Selection for Language Models
von: Li, Hongming, et al.
Veröffentlicht: (2026)
von: Li, Hongming, et al.
Veröffentlicht: (2026)
On the Worst Prompt Performance of Large Language Models
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
Reasons to Reject? Aligning Language Models with Judgments
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
DepWiGNN: A Depth-wise Graph Neural Network for Multi-hop Spatial Reasoning in Text
von: Li, Shuaiyi, et al.
Veröffentlicht: (2023)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2023)
FSTA-SNN:Frequency-based Spatial-Temporal Attention Module for Spiking Neural Networks
von: Yu, Kairong, et al.
Veröffentlicht: (2024)
von: Yu, Kairong, et al.
Veröffentlicht: (2024)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2024) -
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
von: Deng, Chenlong, et al.
Veröffentlicht: (2025) -
InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
von: Li, Shuaiyi, et al.
Veröffentlicht: (2025) -
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
von: Li, Shuaiyi, et al.
Veröffentlicht: (2026) -
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
von: Ma, Junyu, et al.
Veröffentlicht: (2025)