Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Liliang, Chen, Congcong, Xu, Haoran, Kim, Young Jin, Atkinson, Adam, Zhan, Zheng, Sun, Jiankai, Peng, Baolin, Liu, Liyuan, Wang, Shuohang, Cheng, Hao, Gao, Jianfeng, Chen, Weizhu, Shen, Yelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
von: Ren, Liliang, et al.
Veröffentlicht: (2024)
von: Ren, Liliang, et al.
Veröffentlicht: (2024)
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
von: Ren, Liliang, et al.
Veröffentlicht: (2026)
von: Ren, Liliang, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
von: Li, Zichong, et al.
Veröffentlicht: (2026)
von: Li, Zichong, et al.
Veröffentlicht: (2026)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
Test-time Recursive Thinking: Self-Improvement without External Feedback
von: Zhuang, Yufan, et al.
Veröffentlicht: (2026)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2026)
ThetaEvolve: Test-time Learning on Open Problems
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
von: Ren, Bo, et al.
Veröffentlicht: (2026)
von: Ren, Bo, et al.
Veröffentlicht: (2026)
Multi-LoRA Composition for Image Generation
von: Zhong, Ming, et al.
Veröffentlicht: (2024)
von: Zhong, Ming, et al.
Veröffentlicht: (2024)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
GRIN: GRadient-INformed MoE
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
StreamAdapter: Efficient Test Time Adaptation from Contextual Streams
von: Muhtar, Dilxat, et al.
Veröffentlicht: (2024)
von: Muhtar, Dilxat, et al.
Veröffentlicht: (2024)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
von: Das, Souvik, et al.
Veröffentlicht: (2024)
von: Das, Souvik, et al.
Veröffentlicht: (2024)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
von: Ge, Tao, et al.
Veröffentlicht: (2026)
von: Ge, Tao, et al.
Veröffentlicht: (2026)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
von: Wang, Yiping, et al.
Veröffentlicht: (2024)
von: Wang, Yiping, et al.
Veröffentlicht: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
SAS: Simulated Attention Score
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning
von: Yang, Yaming, et al.
Veröffentlicht: (2024)
von: Yang, Yaming, et al.
Veröffentlicht: (2024)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
Architectural Exploration of Hybrid Neural Decoders for Neuromorphic Implantable BMI
von: Mohan, Vivek, et al.
Veröffentlicht: (2025)
von: Mohan, Vivek, et al.
Veröffentlicht: (2025)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
von: Chen, Ting, et al.
Veröffentlicht: (2026)
von: Chen, Ting, et al.
Veröffentlicht: (2026)
SciAgent: Tool-augmented Language Models for Scientific Reasoning
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
von: Ren, Peng, et al.
Veröffentlicht: (2025)
von: Ren, Peng, et al.
Veröffentlicht: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
von: Xu, Haoran, et al.
Veröffentlicht: (2025) -
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
von: Ren, Liliang, et al.
Veröffentlicht: (2024) -
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
von: Ren, Liliang, et al.
Veröffentlicht: (2026) -
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
von: Wang, Yiping, et al.
Veröffentlicht: (2025) -
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
von: Li, Zichong, et al.
Veröffentlicht: (2026)