SplitReason: Learning To Offload Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Akhauri, Yash, Fei, Anthony, Chang, Chi-Chih, AbouElhamayed, Ahmed F., Li, Yueying, Abdelfattah, Mohamed S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
by: Dotzel, Jordan, et al.
Published: (2024)
by: Dotzel, Jordan, et al.
Published: (2024)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Compute Where it Counts: Self Optimizing Language Models
by: Akhauri, Yash, et al.
Published: (2026)
by: Akhauri, Yash, et al.
Published: (2026)
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
by: AbouElhamayed, Ahmed F., et al.
Published: (2024)
by: AbouElhamayed, Ahmed F., et al.
Published: (2024)
Attamba: Attending To Multi-Token States
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
Encodings for Prediction-based Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
On Latency Predictors for Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
NITRO: LLM Inference on Intel Laptop NPUs
by: Fei, Anthony, et al.
Published: (2024)
by: Fei, Anthony, et al.
Published: (2024)
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
by: Hu, Zhanqiu, et al.
Published: (2025)
by: Hu, Zhanqiu, et al.
Published: (2025)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
by: Wang, Pei-Shuo, et al.
Published: (2025)
by: Wang, Pei-Shuo, et al.
Published: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
by: Chang, Chi-Chih, et al.
Published: (2025)
by: Chang, Chi-Chih, et al.
Published: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
by: Chen, Yuzong, et al.
Published: (2024)
by: Chen, Yuzong, et al.
Published: (2024)
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
by: Agarwal, Aradhye, et al.
Published: (2026)
by: Agarwal, Aradhye, et al.
Published: (2026)
Faster LLM Inference via Sequential Monte Carlo
by: Emara, Yahya, et al.
Published: (2026)
by: Emara, Yahya, et al.
Published: (2026)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
by: Ahmed, Seif, et al.
Published: (2025)
by: Ahmed, Seif, et al.
Published: (2025)
Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication
by: Sójka, Stanisław, et al.
Published: (2026)
by: Sójka, Stanisław, et al.
Published: (2026)
Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization
by: Gui, Runquan, et al.
Published: (2026)
by: Gui, Runquan, et al.
Published: (2026)
RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
by: Chen, Yongchao, et al.
Published: (2025)
by: Chen, Yongchao, et al.
Published: (2025)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
by: Sadana, Ananya, et al.
Published: (2025)
by: Sadana, Ananya, et al.
Published: (2025)
Reward Reasoning Model
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
by: Patil, Parth, et al.
Published: (2026)
by: Patil, Parth, et al.
Published: (2026)
Entropy-Guided Reasoning Compression
by: Zhu, Hourun, et al.
Published: (2025)
by: Zhu, Hourun, et al.
Published: (2025)
Multilingual Reasoning Gym: Multilingual Scaling of Procedural Reasoning Environments
by: Dobler, Konstantin, et al.
Published: (2026)
by: Dobler, Konstantin, et al.
Published: (2026)
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
Long Is More Important Than Difficult for Training Reasoning Models
by: Shen, Si, et al.
Published: (2025)
by: Shen, Si, et al.
Published: (2025)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
From Roots to Rewards: Dynamic Tree Reasoning with Reinforcement Learning
by: Bahloul, Ahmed, et al.
Published: (2025)
by: Bahloul, Ahmed, et al.
Published: (2025)
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025)
by: Rameshkumar, Revanth, et al.
Published: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Learning to Reason via Mixture-of-Thought for Logical Reasoning
by: Zheng, Tong, et al.
Published: (2025)
by: Zheng, Tong, et al.
Published: (2025)
Learning to Reason Across Parallel Samples for LLM Reasoning
by: Qi, Jianing, et al.
Published: (2025)
by: Qi, Jianing, et al.
Published: (2025)
Similar Items
-
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
by: Dotzel, Jordan, et al.
Published: (2024) -
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024) -
Compute Where it Counts: Self Optimizing Language Models
by: Akhauri, Yash, et al.
Published: (2026) -
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
by: AbouElhamayed, Ahmed F., et al.
Published: (2024)