Test-Time Personalization: A Diagnostic Framework and Probabilistic Fix for Scaling Failures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Linhai, He, Yulan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probabilistic Test-Time Generalization by Variational Neighbor-Labeling
von: Ambekar, Sameer, et al.
Veröffentlicht: (2023)
von: Ambekar, Sameer, et al.
Veröffentlicht: (2023)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
von: Miyamoto, Sora, et al.
Veröffentlicht: (2026)
von: Miyamoto, Sora, et al.
Veröffentlicht: (2026)
T-POP: Test-Time Personalization with Online Preference Feedback
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
S*: Test Time Scaling for Code Generation
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
Code Generation by Differential Test Time Scaling
von: He, Yifeng, et al.
Veröffentlicht: (2026)
von: He, Yifeng, et al.
Veröffentlicht: (2026)
$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models
von: Bilal, Ahsan, et al.
Veröffentlicht: (2026)
von: Bilal, Ahsan, et al.
Veröffentlicht: (2026)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Scaling Up Probabilistic Circuits by Latent Variable Distillation
von: Liu, Anji, et al.
Veröffentlicht: (2022)
von: Liu, Anji, et al.
Veröffentlicht: (2022)
Parametric Prior Mapping Framework for Non-stationary Probabilistic Time Series Forecasting
von: Li, Jinglin, et al.
Veröffentlicht: (2026)
von: Li, Jinglin, et al.
Veröffentlicht: (2026)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Exploring the Potential of Probabilistic Transformer for Time Series Modeling: A Report on the ST-PT Framework
von: Xiong, Zhangzhi, et al.
Veröffentlicht: (2026)
von: Xiong, Zhangzhi, et al.
Veröffentlicht: (2026)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
Cross-Sample Augmented Test-Time Adaptation for Personalized Intraoperative Hypotension Prediction
von: Li, Kanxue, et al.
Veröffentlicht: (2025)
von: Li, Kanxue, et al.
Veröffentlicht: (2025)
PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency
von: Liu, Zhangyi, et al.
Veröffentlicht: (2026)
von: Liu, Zhangyi, et al.
Veröffentlicht: (2026)
Probabilistic Federated Learning on Uncertain and Heterogeneous Data with Model Personalization
von: Rahman, Ratun, et al.
Veröffentlicht: (2026)
von: Rahman, Ratun, et al.
Veröffentlicht: (2026)
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
von: Tran, Dao, et al.
Veröffentlicht: (2026)
von: Tran, Dao, et al.
Veröffentlicht: (2026)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
Core-Halo Decomposition: Decentralizing Large-Scale Fixed-Point Problems
von: Haixiang, et al.
Veröffentlicht: (2026)
von: Haixiang, et al.
Veröffentlicht: (2026)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
von: Pal, Arka, et al.
Veröffentlicht: (2024)
von: Pal, Arka, et al.
Veröffentlicht: (2024)
FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
Crosslingual Reasoning through Test-Time Scaling
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
von: Yu, Chao, et al.
Veröffentlicht: (2025)
von: Yu, Chao, et al.
Veröffentlicht: (2025)
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
A Pretrained Probabilistic Transformer for City-Scale Traffic Volume Prediction
von: Shen, Shiyu, et al.
Veröffentlicht: (2025)
von: Shen, Shiyu, et al.
Veröffentlicht: (2025)
FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
von: Xu, Jiajun, et al.
Veröffentlicht: (2026)
von: Xu, Jiajun, et al.
Veröffentlicht: (2026)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
Atom of Thoughts for Markov LLM Test-Time Scaling
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
A Unified Framework for Human-Allied Learning of Probabilistic Circuits
von: Karanam, Athresh, et al.
Veröffentlicht: (2024)
von: Karanam, Athresh, et al.
Veröffentlicht: (2024)
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Probabilistic Test-Time Generalization by Variational Neighbor-Labeling
von: Ambekar, Sameer, et al.
Veröffentlicht: (2023) -
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
von: Miyamoto, Sora, et al.
Veröffentlicht: (2026) -
T-POP: Test-Time Personalization with Online Preference Feedback
von: Qu, Zikun, et al.
Veröffentlicht: (2025) -
S*: Test Time Scaling for Code Generation
von: Li, Dacheng, et al.
Veröffentlicht: (2025) -
Code Generation by Differential Test Time Scaling
von: He, Yifeng, et al.
Veröffentlicht: (2026)