SA-DiffuSeq: Addressing Computational and Scalability Challenges in Long-Document Generation with Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Christoforos, Alexandros, Davis, Chadbourne |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE-DiffuSeq: Enhancing Long-Document Diffusion Models with Sparse Attention and Mixture of Experts
by: Christoforos, Alexandros, et al.
Published: (2025)
by: Christoforos, Alexandros, et al.
Published: (2025)
LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification
by: Driscoll, Rory, et al.
Published: (2026)
by: Driscoll, Rory, et al.
Published: (2026)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
by: Li, Zherui, et al.
Published: (2025)
by: Li, Zherui, et al.
Published: (2025)
DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis
by: Long, Shu, et al.
Published: (2026)
by: Long, Shu, et al.
Published: (2026)
Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
by: Hu, Xiang, et al.
Published: (2025)
by: Hu, Xiang, et al.
Published: (2025)
Exploiting the Potential of Seq2Seq Models as Robust Few-Shot Learners
by: Lee, Jihyeon, et al.
Published: (2023)
by: Lee, Jihyeon, et al.
Published: (2023)
Multi-Stage Balanced Distillation: Addressing Long-Tail Challenges in Sequence-Level Knowledge Distillation
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Detection of Suicidal Risk on Social Media: A Hybrid Model
by: Yang, Zaihan, et al.
Published: (2025)
by: Yang, Zaihan, et al.
Published: (2025)
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
by: Phuc, Luu Huu, et al.
Published: (2026)
by: Phuc, Luu Huu, et al.
Published: (2026)
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets
by: Joshi, Harshit, et al.
Published: (2026)
by: Joshi, Harshit, et al.
Published: (2026)
Reference-Free Reinforcement Learning Fine-Tuning for MT: A Seq2Seq Perspective
by: Garcia-Estrada, Ernesto, et al.
Published: (2026)
by: Garcia-Estrada, Ernesto, et al.
Published: (2026)
Undesirable Biases in NLP: Addressing Challenges of Measurement
by: van der Wal, Oskar, et al.
Published: (2022)
by: van der Wal, Oskar, et al.
Published: (2022)
Seq2Seq2Seq: Lossless Data Compression via Discrete Latent Transformers and Reinforcement Learning
by: Khodabandeh, Mahdi, et al.
Published: (2026)
by: Khodabandeh, Mahdi, et al.
Published: (2026)
Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla
by: Das, Sourav Kumar, et al.
Published: (2025)
by: Das, Sourav Kumar, et al.
Published: (2025)
Addressing Longstanding Challenges in Cognitive Science with Language Models
by: Wulff, Dirk U., et al.
Published: (2025)
by: Wulff, Dirk U., et al.
Published: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
SA-MDKIF: A Scalable and Adaptable Medical Domain Knowledge Injection Framework for Large Language Models
by: Xu, Tianhan, et al.
Published: (2024)
by: Xu, Tianhan, et al.
Published: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
ReflectDiffu:Reflect between Emotion-intent Contagion and Mimicry for Empathetic Response Generation via a RL-Diffusion Framework
by: Yuan, Jiahao, et al.
Published: (2024)
by: Yuan, Jiahao, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
by: Chuang, Yun-Yen, et al.
Published: (2024)
by: Chuang, Yun-Yen, et al.
Published: (2024)
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
by: Gao, Yizhao, et al.
Published: (2026)
by: Gao, Yizhao, et al.
Published: (2026)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
by: Lou, Yuxuan, et al.
Published: (2026)
by: Lou, Yuxuan, et al.
Published: (2026)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
Multi-Agent Interactive Question Generation Framework for Long Document Understanding
by: Wang, Kesen, et al.
Published: (2025)
by: Wang, Kesen, et al.
Published: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
Automatic Summarization of Long Documents
by: Chhibbar, Naman, et al.
Published: (2024)
by: Chhibbar, Naman, et al.
Published: (2024)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
by: Hu, Junhao, et al.
Published: (2026)
by: Hu, Junhao, et al.
Published: (2026)
Similar Items
-
MoE-DiffuSeq: Enhancing Long-Document Diffusion Models with Sparse Attention and Mixture of Experts
by: Christoforos, Alexandros, et al.
Published: (2025) -
LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification
by: Driscoll, Rory, et al.
Published: (2026) -
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025) -
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
by: Li, Guanghao, et al.
Published: (2025) -
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
by: Liu, Dong, et al.
Published: (2025)