Similar Items
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
by: Song, Chendong, et al.
Published: (2026)
by: Song, Chendong, et al.
Published: (2026)
UMoE: Unifying Attention and FFN with Shared Experts
by: Yang, Yuanhang, et al.
Published: (2025)
by: Yang, Yuanhang, et al.
Published: (2025)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
Reasoning: From Reflection to Solution
by: Li, Zixi
Published: (2025)
by: Li, Zixi
Published: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
by: Zhou, Hanzhang, et al.
Published: (2024)
by: Zhou, Hanzhang, et al.
Published: (2024)
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
by: Xu, Yao, et al.
Published: (2025)
by: Xu, Yao, et al.
Published: (2025)
Asterisk Operator
by: Li, Zixi
Published: (2025)
by: Li, Zixi
Published: (2025)
Trained Persistent Memory for Frozen Encoder--Decoder LLMs: Six Architectural Methods
by: Jeong, Hong
Published: (2026)
by: Jeong, Hong
Published: (2026)
Decoders Laugh as Loud as Encoders
by: Borodach, Eli, et al.
Published: (2025)
by: Borodach, Eli, et al.
Published: (2025)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
by: Gautam, Aayush, et al.
Published: (2026)
by: Gautam, Aayush, et al.
Published: (2026)
SEED: A Structural Encoder for Embedding-Driven Decoding in Time Series Prediction with LLMs
by: Li, Fengze, et al.
Published: (2025)
by: Li, Fengze, et al.
Published: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
Spatio-Temporal Forecasting of PM2.5 via Spatial-Diffusion guided Encoder-Decoder Architecture
by: Pandey, Malay, et al.
Published: (2024)
by: Pandey, Malay, et al.
Published: (2024)
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits
by: Dikshit, Subrit, et al.
Published: (2025)
by: Dikshit, Subrit, et al.
Published: (2025)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
by: Zhang, Zheng
Published: (2025)
by: Zhang, Zheng
Published: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
On Linearizing Structured Data in Encoder-Decoder Language Models: Insights from Text-to-SQL
by: Shao, Yutong, et al.
Published: (2024)
by: Shao, Yutong, et al.
Published: (2024)
An Attention Mechanism for Robust Multimodal Integration in a Global Workspace Architecture
by: Bertin-Johannet, Roland, et al.
Published: (2026)
by: Bertin-Johannet, Roland, et al.
Published: (2026)
Cross-Attention and Encoder-Decoder Transformers: A Logical Characterization
by: Ahvonen, Veeti, et al.
Published: (2026)
by: Ahvonen, Veeti, et al.
Published: (2026)
FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus
by: Fu, Yuxin, et al.
Published: (2024)
by: Fu, Yuxin, et al.
Published: (2024)
Multi-task Federated Learning with Encoder-Decoder Structure: Enabling Collaborative Learning Across Different Tasks
by: Zhou, Jingxuan, et al.
Published: (2025)
by: Zhou, Jingxuan, et al.
Published: (2025)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
by: Zhao, Yanjun, et al.
Published: (2024)
by: Zhao, Yanjun, et al.
Published: (2024)
Jewelry Recognition via Encoder-Decoder Models
by: Alcalde-Llergo, José M., et al.
Published: (2024)
by: Alcalde-Llergo, José M., et al.
Published: (2024)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
by: Arriola, Marianne, et al.
Published: (2025)
by: Arriola, Marianne, et al.
Published: (2025)
PACAD-Based Structural Reasoning for AGI: From Probabilistic Pattern-Matching to Verifiable Canonical Architectures – paper 1, version 2
by: Brown, Cameron
Published: (2026)
by: Brown, Cameron
Published: (2026)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
by: Liu, Ningyuan, et al.
Published: (2025)
by: Liu, Ningyuan, et al.
Published: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026)
by: Wu, Hanjiang, et al.
Published: (2026)
Chain-in-Tree: Back to Sequential Reasoning in LLM Tree Search
by: Li, Xinzhe
Published: (2025)
by: Li, Xinzhe
Published: (2025)
PACAD-Based Structural Reasoning for AGI: From Probabilistic Pattern-Matching to Verifiable Canonical Architectures – paper 1, version 1
by: Brown, Cameron
Published: (2026)
by: Brown, Cameron
Published: (2026)
A Novel Architecture for Symbolic Reasoning with Decision Trees and LLM Agents
by: Kiruluta, Andrew
Published: (2025)
by: Kiruluta, Andrew
Published: (2025)
TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series Forecasting
by: Lee, Jaebin, et al.
Published: (2025)
by: Lee, Jaebin, et al.
Published: (2025)
Towards A Universal Graph Structural Encoder
by: Chen, Jialin, et al.
Published: (2025)
by: Chen, Jialin, et al.
Published: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
by: Xu, Zicheng, et al.
Published: (2025)
by: Xu, Zicheng, et al.
Published: (2025)
Cross-Attention Speculative Decoding
by: Zhong, Wei, et al.
Published: (2025)
by: Zhong, Wei, et al.
Published: (2025)
LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories
by: Kang, Liwei, et al.
Published: (2026)
by: Kang, Liwei, et al.
Published: (2026)
Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence Tree
by: Peng, Qi, et al.
Published: (2025)
by: Peng, Qi, et al.
Published: (2025)
SIEDD: Shared-Implicit Encoder with Discrete Decoders
by: Rangarajan, Vikram, et al.
Published: (2025)
by: Rangarajan, Vikram, et al.
Published: (2025)
Similar Items
-
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
by: Song, Chendong, et al.
Published: (2026) -
UMoE: Unifying Attention and FFN with Shared Experts
by: Yang, Yuanhang, et al.
Published: (2025) -
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026) -
Reasoning: From Reflection to Solution
by: Li, Zixi
Published: (2025) -
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)