Gespeichert in:
| Hauptverfasser: | Zuo, Bowen, Zhou, Dongruo, Zhu, Yinglun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.21018 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
Active Testing of Large Language Models via Approximate Neyman Allocation
von: Liu, Zeli, et al.
Veröffentlicht: (2026)
von: Liu, Zeli, et al.
Veröffentlicht: (2026)
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
von: Zhao, Runze, et al.
Veröffentlicht: (2025)
von: Zhao, Runze, et al.
Veröffentlicht: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
Interactive Machine Learning: From Theory to Scale
von: Zhu, Yinglun
Veröffentlicht: (2025)
von: Zhu, Yinglun
Veröffentlicht: (2025)
Adaptive Test-Time Compute Allocation via Learned Heuristics over Categorical Structure
von: Qu, Shuhui
Veröffentlicht: (2026)
von: Qu, Shuhui
Veröffentlicht: (2026)
Online Finetuning Decision Transformers with Pure RL Gradients
von: Luo, Junkai, et al.
Veröffentlicht: (2026)
von: Luo, Junkai, et al.
Veröffentlicht: (2026)
Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data
von: Zhang, Jiancheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiancheng, et al.
Veröffentlicht: (2025)
Mixtraining: A Better Trade-Off Between Compute and Performance
von: Li, Zexin, et al.
Veröffentlicht: (2025)
von: Li, Zexin, et al.
Veröffentlicht: (2025)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
von: Yang, Chen, et al.
Veröffentlicht: (2024)
von: Yang, Chen, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
von: Zhang, Junkai, et al.
Veröffentlicht: (2024)
von: Zhang, Junkai, et al.
Veröffentlicht: (2024)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
von: Paglieri, Davide, et al.
Veröffentlicht: (2025)
von: Paglieri, Davide, et al.
Veröffentlicht: (2025)
Privacy-Aware RAG: Secure and Isolated Knowledge Retrieval
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
Provably Secure Retrieval-Augmented Generation
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
Provable Zero-Shot Generalization in Offline Reinforcement Learning
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
Change of Thought: Adaptive Test-Time Computation
von: Mathur, Mrinal, et al.
Veröffentlicht: (2025)
von: Mathur, Mrinal, et al.
Veröffentlicht: (2025)
Why do AI agents communicate in human language?
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
Adaptive Rectification Sampling for Test-Time Compute Scaling
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
How to Provably Improve Return Conditioned Supervised Learning?
von: Liu, Zhishuai, et al.
Veröffentlicht: (2025)
von: Liu, Zhishuai, et al.
Veröffentlicht: (2025)
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Zero-Overhead Introspection for Adaptive Test-Time Compute
von: Manvi, Rohin, et al.
Veröffentlicht: (2025)
von: Manvi, Rohin, et al.
Veröffentlicht: (2025)
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
von: Che, Lirong, et al.
Veröffentlicht: (2026)
von: Che, Lirong, et al.
Veröffentlicht: (2026)
DSevolve: Enabling Real-Time Adaptive Scheduling on Dynamic Shop Floor with LLM-Evolved Heuristic Portfolios
von: Huang, Jin, et al.
Veröffentlicht: (2026)
von: Huang, Jin, et al.
Veröffentlicht: (2026)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
von: Islam, Md Mirajul, et al.
Veröffentlicht: (2026)
von: Islam, Md Mirajul, et al.
Veröffentlicht: (2026)
Real-Time Reasoning Agents in Evolving Environments
von: Wen, Yule, et al.
Veröffentlicht: (2025)
von: Wen, Yule, et al.
Veröffentlicht: (2025)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
von: Jin, Chen, et al.
Veröffentlicht: (2026)
von: Jin, Chen, et al.
Veröffentlicht: (2026)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
Discover Fast Power Allocation Solution for Multi-Target Tracking via AlphaEvolve Evolution
von: Hou, Zhenkang, et al.
Veröffentlicht: (2026)
von: Hou, Zhenkang, et al.
Veröffentlicht: (2026)
In-Context Learning with Iterative Demonstration Selection
von: Qin, Chengwei, et al.
Veröffentlicht: (2023)
von: Qin, Chengwei, et al.
Veröffentlicht: (2023)
Rectifying Demonstration Shortcut in In-Context Learning
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
Dynamic Demonstrations Controller for In-Context Learning
von: Zhao, Fei, et al.
Veröffentlicht: (2023)
von: Zhao, Fei, et al.
Veröffentlicht: (2023)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
von: Alomrani, Mohammad Ali, et al.
Veröffentlicht: (2025)
von: Alomrani, Mohammad Ali, et al.
Veröffentlicht: (2025)
Comparable Demonstrations are Important in In-Context Learning: A Novel Perspective on Demonstration Selection
von: Fan, Caoyun, et al.
Veröffentlicht: (2023)
von: Fan, Caoyun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
von: Zuo, Bowen, et al.
Veröffentlicht: (2025) -
Active Testing of Large Language Models via Approximate Neyman Allocation
von: Liu, Zeli, et al.
Veröffentlicht: (2026) -
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
von: Zhao, Runze, et al.
Veröffentlicht: (2025) -
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025) -
Interactive Machine Learning: From Theory to Scale
von: Zhu, Yinglun
Veröffentlicht: (2025)