$V_0$: A Generalist Value Model for Any Policy at State Zero
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yi-Kai, Yao, Zhiyuan, Hao, Hongyan, Sun, Yueqing, Gu, Qi, Su, Hui, Cai, Xunliang, Zhan, De-Chuan, Ye, Han-Jia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
von: Tu, Dunwei, et al.
Veröffentlicht: (2026)
von: Tu, Dunwei, et al.
Veröffentlicht: (2026)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Any-point Trajectory Modeling for Policy Learning
von: Wen, Chuan, et al.
Veröffentlicht: (2023)
von: Wen, Chuan, et al.
Veröffentlicht: (2023)
TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training
von: Yang, Jinluan, et al.
Veröffentlicht: (2026)
von: Yang, Jinluan, et al.
Veröffentlicht: (2026)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
One-Embedding-Fits-All: Efficient Zero-Shot Time Series Forecasting by a Model Zoo
von: Shi, Hao-Nan, et al.
Veröffentlicht: (2025)
von: Shi, Hao-Nan, et al.
Veröffentlicht: (2025)
A Closer Look at Deep Learning Methods on Tabular Datasets
von: Ye, Han-Jia, et al.
Veröffentlicht: (2024)
von: Ye, Han-Jia, et al.
Veröffentlicht: (2024)
Unraveling the Mystery of Scaling Laws: Part I
von: Su, Hui, et al.
Veröffentlicht: (2024)
von: Su, Hui, et al.
Veröffentlicht: (2024)
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement
von: Hu, Zhaofeng, et al.
Veröffentlicht: (2026)
von: Hu, Zhaofeng, et al.
Veröffentlicht: (2026)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens
von: Huang, Ting-Ji, et al.
Veröffentlicht: (2024)
von: Huang, Ting-Ji, et al.
Veröffentlicht: (2024)
Adaptive Adapter Routing for Long-Tailed Class-Incremental Learning
von: Qi, Zhi-Hong, et al.
Veröffentlicht: (2024)
von: Qi, Zhi-Hong, et al.
Veröffentlicht: (2024)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2025)
TV100: A TV Series Dataset that Pre-Trained CLIP Has Not Seen
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
Enhancement of Photo‐ and Electrocatalytic H 2 O 2 Production by the Reversible Catechol‐to‐ o ‐Benzoquinone Transformation in Polydopamine
von: Yueqing Jia, et al.
Veröffentlicht: (2025)
von: Yueqing Jia, et al.
Veröffentlicht: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Taxonomy-Guided Zero-Shot Recommendations with LLMs
von: Liang, Yueqing, et al.
Veröffentlicht: (2024)
von: Liang, Yueqing, et al.
Veröffentlicht: (2024)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
von: Yi, Chao, et al.
Veröffentlicht: (2024)
von: Yi, Chao, et al.
Veröffentlicht: (2024)
Learning to Self-Verify Makes Language Models Better Reasoners
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning
von: Zhu, Jiapeng, et al.
Veröffentlicht: (2026)
von: Zhu, Jiapeng, et al.
Veröffentlicht: (2026)
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
von: Jing, Junpeng, et al.
Veröffentlicht: (2025)
von: Jing, Junpeng, et al.
Veröffentlicht: (2025)
Possible Value Analysis based on Symbolic Lattice
von: Zhan, Qi
Veröffentlicht: (2024)
von: Zhan, Qi
Veröffentlicht: (2024)
Model Assembly Learning with Heterogeneous Layer Weight Merging
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
Revisiting Class-Incremental Learning with Pre-Trained Models: Generalizability and Adaptivity are All You Need
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2023)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2023)
Dual Consolidation for Pre-Trained Model-Based Domain-Incremental Learning
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
AnyI2V: Animating Any Conditional Image with Motion Control
von: Li, Ziye, et al.
Veröffentlicht: (2025)
von: Li, Ziye, et al.
Veröffentlicht: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Generalist predators function as pest specialists: Examining diet composition of spiders and ladybeetles across rice crop stages
von: Gen‐Chang Hsu, et al.
Veröffentlicht: (2025)
von: Gen‐Chang Hsu, et al.
Veröffentlicht: (2025)
Submonthly timescale oscillation characteristics of the East Asian winter monsoon and its effect on the temperature of southwest China in 2010.
von: Qi, Dongmei, et al.
Veröffentlicht: (2016)
von: Qi, Dongmei, et al.
Veröffentlicht: (2016)
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026) -
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
von: Tu, Dunwei, et al.
Veröffentlicht: (2026) -
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026) -
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025) -
Any-point Trajectory Modeling for Policy Learning
von: Wen, Chuan, et al.
Veröffentlicht: (2023)