FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hao Mark, Mo, Zhiwen, Lu, Guanxi, Liang, Shuang, Ma, Lingxiao, Luk, Wayne, Fan, Hongxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error Correction
by: Campbell, Charlie, et al.
Published: (2025)
by: Campbell, Charlie, et al.
Published: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
by: Mo, Zhiwen, et al.
Published: (2026)
by: Mo, Zhiwen, et al.
Published: (2026)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
Accelerating MRI Uncertainty Estimation with Mask-based Bayesian Neural Network
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
Context Memorization for Efficient Long Context Generation
by: Okoshi, Yasuyuki, et al.
Published: (2026)
by: Okoshi, Yasuyuki, et al.
Published: (2026)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
by: Chen, Hao Mark, et al.
Published: (2026)
by: Chen, Hao Mark, et al.
Published: (2026)
Versatile Cross-platform Compilation Toolchain for Schrödinger-style Quantum Circuit Simulation
by: Lu, Yuncheng, et al.
Published: (2025)
by: Lu, Yuncheng, et al.
Published: (2025)
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
by: Que, Zhiqiang, et al.
Published: (2025)
by: Que, Zhiqiang, et al.
Published: (2025)
Enhancing Dropout-based Bayesian Neural Networks with Multi-Exit on FPGA
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
by: Chen, Jie, et al.
Published: (2025)
by: Chen, Jie, et al.
Published: (2025)
The Adjoint Polynomial of a Polyhedral Cone is a Covolume Polynomial
by: Li, Guanxi
Published: (2025)
by: Li, Guanxi
Published: (2025)
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
by: Shen, Guanxi
Published: (2025)
by: Shen, Guanxi
Published: (2025)
Segre Class of Schemes with Regularly Embedded Components
by: Li, Guanxi
Published: (2025)
by: Li, Guanxi
Published: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Accelerating 3D Gaussian Splatting with Neural Sorting and Axis-Oriented Rasterization
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
by: Chen, Zhekai, et al.
Published: (2025)
by: Chen, Zhekai, et al.
Published: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Scalable Time-Series Causal Discovery with Approximate Causal Ordering
by: Jiao, Ziyang, et al.
Published: (2024)
by: Jiao, Ziyang, et al.
Published: (2024)
Robust Time Series Causal Discovery for Agent-Based Model Validation
by: Yu, Gene, et al.
Published: (2024)
by: Yu, Gene, et al.
Published: (2024)
VCDF: A Validated Consensus-Driven Framework for Time Series Causal Discovery
by: Yu, Gene, et al.
Published: (2026)
by: Yu, Gene, et al.
Published: (2026)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
by: Zhang, Zehuan, et al.
Published: (2026)
by: Zhang, Zehuan, et al.
Published: (2026)
Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines
by: Papaphilippou, Philippos, et al.
Published: (2024)
by: Papaphilippou, Philippos, et al.
Published: (2024)
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
by: Zhou, Shu, et al.
Published: (2026)
by: Zhou, Shu, et al.
Published: (2026)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
by: Que, Zhiqiang, et al.
Published: (2022)
by: Que, Zhiqiang, et al.
Published: (2022)
Testing Gravity with Realistic Gravitational Waveforms in Pulsar Timing Arrays
by: Hu, Wayne, et al.
Published: (2024)
by: Hu, Wayne, et al.
Published: (2024)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
by: Wu, Haoran, et al.
Published: (2025)
by: Wu, Haoran, et al.
Published: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
by: Liu, Zhijun, et al.
Published: (2024)
by: Liu, Zhijun, et al.
Published: (2024)
Interpretable Prediction of Late‐Stage CKM Syndrome Association From Dietary Nutrients in Accelerated Aging Using SHAP and LIME
by: Hongxiang Tu, et al.
Published: (2026)
by: Hongxiang Tu, et al.
Published: (2026)
Ranking Reasoning LLMs under Test-Time Scaling
by: Hariri, Mohsen, et al.
Published: (2026)
by: Hariri, Mohsen, et al.
Published: (2026)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024)
by: Bi, Zhenni, et al.
Published: (2024)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026)
by: Yu, Yongcan, et al.
Published: (2026)
Similar Items
-
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025) -
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
by: Lu, Guanxi, et al.
Published: (2025) -
Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error Correction
by: Campbell, Charlie, et al.
Published: (2025) -
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
by: Mo, Zhiwen, et al.
Published: (2026) -
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)