Scaling Bidirectional Spans and Span Violations in Attention Mechanism
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jongwook, Yun, Sangheon, Yoon, Sukjin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hallucinated Span Detection with Multi-View Attention Features
by: Ogasa, Yuya, et al.
Published: (2025)
by: Ogasa, Yuya, et al.
Published: (2025)
Learning to Reason for Hallucination Span Detection
by: Su, Hsuan, et al.
Published: (2025)
by: Su, Hsuan, et al.
Published: (2025)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
by: Lyu, Boxuan, et al.
Published: (2025)
by: Lyu, Boxuan, et al.
Published: (2025)
Limitations of Normalization in Attention Mechanism
by: Mudarisov, Timur, et al.
Published: (2025)
by: Mudarisov, Timur, et al.
Published: (2025)
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024)
by: Jin, Tian, et al.
Published: (2024)
Scaling Reasoning without Attention
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
by: Chuang, Yung-Sung, et al.
Published: (2024)
by: Chuang, Yung-Sung, et al.
Published: (2024)
SpanGNN: Towards Memory-Efficient Graph Neural Networks via Spanning Subgraph Training
by: Gu, Xizhi, et al.
Published: (2024)
by: Gu, Xizhi, et al.
Published: (2024)
Asynchronous and Segmented Bidirectional Encoding for NMT
by: Yang, Jingpu, et al.
Published: (2024)
by: Yang, Jingpu, et al.
Published: (2024)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
by: Lee, Changhun, et al.
Published: (2025)
by: Lee, Changhun, et al.
Published: (2025)
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
by: Kim, Siun, et al.
Published: (2026)
by: Kim, Siun, et al.
Published: (2026)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
A Simple Attention-Based Mechanism for Bimodal Emotion Classification
by: Elabd, Mazen, et al.
Published: (2024)
by: Elabd, Mazen, et al.
Published: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
The Laplacian Keyboard: Beyond the Linear Span
by: Chandrasekar, Siddarth, et al.
Published: (2026)
by: Chandrasekar, Siddarth, et al.
Published: (2026)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
by: Kim, Junseo, et al.
Published: (2025)
by: Kim, Junseo, et al.
Published: (2025)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
by: Mao, Yujun, et al.
Published: (2024)
by: Mao, Yujun, et al.
Published: (2024)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
by: Tomczak, Nathaniel, et al.
Published: (2025)
by: Tomczak, Nathaniel, et al.
Published: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
by: Zhu, Xudong, et al.
Published: (2025)
by: Zhu, Xudong, et al.
Published: (2025)
MOGAM: A Multimodal Object-oriented Graph Attention Model for Depression Detection
by: Cha, Junyeop, et al.
Published: (2024)
by: Cha, Junyeop, et al.
Published: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
by: Yoon, Hee Suk, et al.
Published: (2025)
by: Yoon, Hee Suk, et al.
Published: (2025)
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
by: Kim, Dahyun, et al.
Published: (2023)
by: Kim, Dahyun, et al.
Published: (2023)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
by: Huang, Ruizhe, et al.
Published: (2025)
by: Huang, Ruizhe, et al.
Published: (2025)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
by: Zou, Minghui, et al.
Published: (2024)
by: Zou, Minghui, et al.
Published: (2024)
Learning to Interpret Weight Differences in Language Models
by: Goel, Avichal, et al.
Published: (2025)
by: Goel, Avichal, et al.
Published: (2025)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026)
by: Flamant, Cedric, et al.
Published: (2026)
SEER: The Span-based Emotion Evidence Retrieval Benchmark
by: Sampath, Aneesha, et al.
Published: (2025)
by: Sampath, Aneesha, et al.
Published: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
by: Yoa, Seungdong, et al.
Published: (2026)
by: Yoa, Seungdong, et al.
Published: (2026)
Detecting Suicidal Ideation in Text with Interpretable Deep Learning: A CNN-BiGRU with Attention Mechanism
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer Approach
by: Chaturvedi, Rochana, et al.
Published: (2025)
by: Chaturvedi, Rochana, et al.
Published: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Similar Items
-
Hallucinated Span Detection with Multi-View Attention Features
by: Ogasa, Yuya, et al.
Published: (2025) -
Learning to Reason for Hallucination Span Detection
by: Su, Hsuan, et al.
Published: (2025) -
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
by: Wang, Chao, et al.
Published: (2026) -
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024) -
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
by: Lyu, Boxuan, et al.
Published: (2025)