Saved in:
| Main Authors: | Jin, Bowen, Collins, TJ, Yu, Donghan, Cemri, Mert, Zhang, Shenao, Li, Mengyu, Tang, Jay, Qin, Tian, Xu, Zhiyang, Lu, Jiarui, Yin, Guoli, Han, Jiawei, Wang, Zirui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.02755 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
COMPASS: Benchmarking Constrained Optimization in LLM Agents
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
Maintaining nature’s balance while building for man's protection (abstract)
by: Bowen, T.J.
Published: (1967)
by: Bowen, T.J.
Published: (1967)
Adaptive TD-Lambda for Cooperative Multi-agent Reinforcement Learning
by: Deng, Yue, et al.
Published: (2026)
by: Deng, Yue, et al.
Published: (2026)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
by: Bai, Hao, et al.
Published: (2024)
by: Bai, Hao, et al.
Published: (2024)
Active Hypothesis Testing under Computational Budgets with Applications to GWAS and LLM
by: Kuang, Qi, et al.
Published: (2025)
by: Kuang, Qi, et al.
Published: (2025)
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
by: Tian, Runchu, et al.
Published: (2025)
by: Tian, Runchu, et al.
Published: (2025)
AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing
by: Ma, Zhiyuan, et al.
Published: (2023)
by: Ma, Zhiyuan, et al.
Published: (2023)
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
by: Lin, Jiuheng, et al.
Published: (2025)
by: Lin, Jiuheng, et al.
Published: (2025)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
Enhancing Toughness of Epoxy Crossing Network: The Role of Branched Reactive Polyethersulfone Ketone
by: Xiaohuan Li, et al.
Published: (2025)
by: Xiaohuan Li, et al.
Published: (2025)
Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
by: Wu, Junlin, et al.
Published: (2025)
by: Wu, Junlin, et al.
Published: (2025)
Expensive Homeomorphism of Convex Bodies
by: Kim, Donghan
Published: (2025)
by: Kim, Donghan
Published: (2025)
Metric Topologies on Multiset Spaces as Topological Monoids and Their Group Completion
by: Kim, Donghan
Published: (2025)
by: Kim, Donghan
Published: (2025)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
by: Cemri, Mert, et al.
Published: (2026)
by: Cemri, Mert, et al.
Published: (2026)
A Budget-Adaptive Allocation Rule for Optimal Computing Budget Allocation
by: Cao, Zirui, et al.
Published: (2023)
by: Cao, Zirui, et al.
Published: (2023)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
LLM App Store Analysis: A Vision and Roadmap
by: Zhao, Yanjie, et al.
Published: (2024)
by: Zhao, Yanjie, et al.
Published: (2024)
On the Foundations of Trustworthy Artificial Intelligence
by: Dunham, TJ
Published: (2026)
by: Dunham, TJ
Published: (2026)
Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative Watermarking
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis
by: Qin, Bowen
Published: (2026)
by: Qin, Bowen
Published: (2026)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Flexible Pressure Sensors Enhanced by 3D‐Printed Microstructures
by: Yuan Jin, et al.
Published: (2025)
by: Yuan Jin, et al.
Published: (2025)
Hybrid Latent Reasoning via Reinforcement Learning
by: Yue, Zhenrui, et al.
Published: (2025)
by: Yue, Zhenrui, et al.
Published: (2025)
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
by: Fang, Yin, et al.
Published: (2025)
by: Fang, Yin, et al.
Published: (2025)
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Belief Aided Navigation using Bayesian Reinforcement Learning for Avoiding Humans in Blind Spots
by: Kim, Jinyeob, et al.
Published: (2024)
by: Kim, Jinyeob, et al.
Published: (2024)
Let the Barbarians In: How AI Can Accelerate Systems Performance Research
by: Cheng, Audrey, et al.
Published: (2025)
by: Cheng, Audrey, et al.
Published: (2025)
FLM-101B: An Open LLM and How to Train It with $100K Budget
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
by: Zheng, Lifan, et al.
Published: (2025)
by: Zheng, Lifan, et al.
Published: (2025)
Aprovechamiento de alimento vivo Culex quinquefasciatus en la dieta del pez cebra Brachidanio rerio (Pisces: Cyprinidae) con énfasis en la reproducción
by: T.J. Olascoaga
Published: (2005)
by: T.J. Olascoaga
Published: (2005)
MYOPIA IN ADOLESCENTS IN MOUNTAINOUS AND FOOTHILL AREAS OF THE ANDIJAN REGION, THE CAUSES OF THE SPREAD
by: Usmanova T.J.
Published: (2026)
by: Usmanova T.J.
Published: (2026)
PRODUCCION DE PLANTAS DE TOMATE Y CHILE APLICANDO PACLOBUTRAZOL AL FOLLAJE
by: ALCARAZ VELAZQUEZ, TJ
Published: (2008)
by: ALCARAZ VELAZQUEZ, TJ
Published: (2008)
Alimentación del cerdo / T. J. Cunha; traductor, Eduardo Zorita Tomillo
by: Cunha, T.J
by: Cunha, T.J
Producción de plantas de tomate y chile aplicando paclobutrazol al follaje
by: TJ Velázquez-Alcaraz
Published: (2008)
by: TJ Velázquez-Alcaraz
Published: (2008)
Similar Items
-
COMPASS: Benchmarking Constrained Optimization in LLM Agents
by: Qin, Tian, et al.
Published: (2025) -
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025) -
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024) -
Maintaining nature’s balance while building for man's protection (abstract)
by: Bowen, T.J.
Published: (1967) -
Adaptive TD-Lambda for Cooperative Multi-agent Reinforcement Learning
by: Deng, Yue, et al.
Published: (2026)