RW-TTT: Batched Serving for Request-Owned Test-Time Training State
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Jian, Kou, Zhizhuo, Tian, Yao, Zhang, Hao, Chen, Han, Han, Sirui, Guo, Yike |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
par: Liu, Fangfu, et autres
Publié: (2026)
par: Liu, Fangfu, et autres
Publié: (2026)
Meta-TTT: A Meta-learning Minimax Framework For Test-Time Training
par: Tao, Chen, et autres
Publié: (2024)
par: Tao, Chen, et autres
Publié: (2024)
Automate Strategy Finding with LLM in Quant Investment
par: Kou, Zhizhuo, et autres
Publié: (2024)
par: Kou, Zhizhuo, et autres
Publié: (2024)
SR-TTT: Surprisal-Aware Residual Test-Time Training
par: P, Swamynathan V
Publié: (2026)
par: P, Swamynathan V
Publié: (2026)
REE-TTT: Highly Adaptive Radar Echo Extrapolation Based on Test-Time Training
par: Di, Xin, et autres
Publié: (2026)
par: Di, Xin, et autres
Publié: (2026)
NC-TTT: A Noise Contrastive Approach for Test-Time Training
par: Osowiechi, David, et autres
Publié: (2024)
par: Osowiechi, David, et autres
Publié: (2024)
ReC-TTT: Contrastive Feature Reconstruction for Test-Time Training
par: Colussi, Marco, et autres
Publié: (2024)
par: Colussi, Marco, et autres
Publié: (2024)
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong
par: Han, Sirui, et autres
Publié: (2025)
par: Han, Sirui, et autres
Publié: (2025)
scFusionTTT: Single-cell transcriptomics and proteomics fusion with Test-Time Training layers
par: Meng, Dian, et autres
Publié: (2024)
par: Meng, Dian, et autres
Publié: (2024)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
par: Gao, Shihong, et autres
Publié: (2025)
par: Gao, Shihong, et autres
Publié: (2025)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
par: Tang, Yinghao, et autres
Publié: (2025)
par: Tang, Yinghao, et autres
Publié: (2025)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
par: Chan, Chi-Min, et autres
Publié: (2025)
par: Chan, Chi-Min, et autres
Publié: (2025)
Unraveling Batch Normalization for Realistic Test-Time Adaptation
par: Su, Zixian, et autres
Publié: (2023)
par: Su, Zixian, et autres
Publié: (2023)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
par: Li, Lujun, et autres
Publié: (2025)
par: Li, Lujun, et autres
Publié: (2025)
NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training
par: Wang, Suli, et autres
Publié: (2025)
par: Wang, Suli, et autres
Publié: (2025)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
par: Gu, Hao, et autres
Publié: (2026)
par: Gu, Hao, et autres
Publié: (2026)
Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training
par: Yang, Xi, et autres
Publié: (2025)
par: Yang, Xi, et autres
Publié: (2025)
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving
par: Kadadekar, Sahil
Publié: (2026)
par: Kadadekar, Sahil
Publié: (2026)
MemFly: On-the-Fly Memory Optimization via Information Bottleneck
par: Zhang, Zhenyuan, et autres
Publié: (2026)
par: Zhang, Zhenyuan, et autres
Publié: (2026)
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
par: Chen, Junyi, et autres
Publié: (2025)
par: Chen, Junyi, et autres
Publié: (2025)
Diversified Batch Selection for Training Acceleration
par: Hong, Feng, et autres
Publié: (2024)
par: Hong, Feng, et autres
Publié: (2024)
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
par: Zhu, Zhenghao, et autres
Publié: (2025)
par: Zhu, Zhenghao, et autres
Publié: (2025)
Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
par: Chen, Zixuan, et autres
Publié: (2026)
par: Chen, Zixuan, et autres
Publié: (2026)
Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification
par: Cao, Chuxue, et autres
Publié: (2026)
par: Cao, Chuxue, et autres
Publié: (2026)
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
par: Wang, Hao, et autres
Publié: (2026)
par: Wang, Hao, et autres
Publié: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
par: Xu, Binxing, et autres
Publié: (2026)
par: Xu, Binxing, et autres
Publié: (2026)
Resilient Practical Test-Time Adaptation: Soft Batch Normalization Alignment and Entropy-driven Memory Bank
par: Zhou, Xingzhi, et autres
Publié: (2024)
par: Zhou, Xingzhi, et autres
Publié: (2024)
Do Not Wait: Learning Re-Ranking Model Without User Feedback At Serving Time in E-Commerce
par: Wang, Yuan, et autres
Publié: (2024)
par: Wang, Yuan, et autres
Publié: (2024)
Test Time Training for Supervised Causal Learning
par: Deng, Zizhen, et autres
Publié: (2026)
par: Deng, Zizhen, et autres
Publié: (2026)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
par: Zhao, Yilong, et autres
Publié: (2024)
par: Zhao, Yilong, et autres
Publié: (2024)
Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
par: Delavande, Julien, et autres
Publié: (2026)
par: Delavande, Julien, et autres
Publié: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
par: Zhang, Wei, et autres
Publié: (2025)
par: Zhang, Wei, et autres
Publié: (2025)
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
par: Chan, Chi-Min, et autres
Publié: (2026)
par: Chan, Chi-Min, et autres
Publié: (2026)
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
par: Deng, Qi, et autres
Publié: (2024)
par: Deng, Qi, et autres
Publié: (2024)
Requests of a Feather Must Flock Together: Batch Size vs. Prefix Homogeneity in LLM Inference
par: Rathi, Saksham, et autres
Publié: (2026)
par: Rathi, Saksham, et autres
Publié: (2026)
Higher-Order Asymptotics of Test-Time Adaptation for Batch Normalization Statistics
par: Kimura, Masanari
Publié: (2025)
par: Kimura, Masanari
Publié: (2025)
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
par: Lin, Chaofan, et autres
Publié: (2024)
par: Lin, Chaofan, et autres
Publié: (2024)
Structural Alignment Improves Graph Test-Time Adaptation
par: Hsu, Hans Hao-Hsun, et autres
Publié: (2025)
par: Hsu, Hans Hao-Hsun, et autres
Publié: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
par: Chen, Lequn, et autres
Publié: (2023)
par: Chen, Lequn, et autres
Publié: (2023)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
par: Akyürek, Ekin, et autres
Publié: (2024)
par: Akyürek, Ekin, et autres
Publié: (2024)
Documents similaires
-
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
par: Liu, Fangfu, et autres
Publié: (2026) -
Meta-TTT: A Meta-learning Minimax Framework For Test-Time Training
par: Tao, Chen, et autres
Publié: (2024) -
Automate Strategy Finding with LLM in Quant Investment
par: Kou, Zhizhuo, et autres
Publié: (2024) -
SR-TTT: Surprisal-Aware Residual Test-Time Training
par: P, Swamynathan V
Publié: (2026) -
REE-TTT: Highly Adaptive Radar Echo Extrapolation Based on Test-Time Training
par: Di, Xin, et autres
Publié: (2026)