Gespeichert in:
| Hauptverfasser: | Wu, JieHao, Wang, Ziwei, Sheng, Junjie, Li, Wenhao, Wang, Xiangfeng, Luo, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.10117 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Reinforcement Learning for Virtual Machine Scheduling
von: Sheng, Junjie, et al.
Veröffentlicht: (2025)
von: Sheng, Junjie, et al.
Veröffentlicht: (2025)
TextAtari: 100K Frames Game Playing with Language Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
von: Wu, Jiehao, et al.
Veröffentlicht: (2026)
von: Wu, Jiehao, et al.
Veröffentlicht: (2026)
Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation
von: Fang, Jingzhi, et al.
Veröffentlicht: (2026)
von: Fang, Jingzhi, et al.
Veröffentlicht: (2026)
Disentangled Representation Learning with Large Language Models for Text-Attributed Graphs
von: Qin, Yijian, et al.
Veröffentlicht: (2023)
von: Qin, Yijian, et al.
Veröffentlicht: (2023)
MiniDisc: Minimal Distillation Schedule for Language Model Compression
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
Promoting Data and Model Privacy in Federated Learning through Quantized LoRA
von: Zhu, JianHao, et al.
Veröffentlicht: (2024)
von: Zhu, JianHao, et al.
Veröffentlicht: (2024)
AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2025)
Mental Modeling of Reinforcement Learning Agents by Language Models
von: Lu, Wenhao, et al.
Veröffentlicht: (2024)
von: Lu, Wenhao, et al.
Veröffentlicht: (2024)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
von: You, Jian, et al.
Veröffentlicht: (2024)
von: You, Jian, et al.
Veröffentlicht: (2024)
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
von: Wang, Fali, et al.
Veröffentlicht: (2025)
von: Wang, Fali, et al.
Veröffentlicht: (2025)
Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Synthetic Sandbox for Training Machine Learning Engineering Agents
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
Adaptive Social Learning via Mode Policy Optimization for Language Agents
von: Wang, Minzheng, et al.
Veröffentlicht: (2025)
von: Wang, Minzheng, et al.
Veröffentlicht: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
von: Li, Zelong, et al.
Veröffentlicht: (2024)
von: Li, Zelong, et al.
Veröffentlicht: (2024)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
GraphThought: Graph Combinatorial Optimization with Thought Generation
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
CodeACT: Code Adaptive Compute-efficient Tuning Framework for Code LLMs
von: Lv, Weijie, et al.
Veröffentlicht: (2024)
von: Lv, Weijie, et al.
Veröffentlicht: (2024)
Direct Multi-Turn Preference Optimization for Language Agents
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
A Survey of Automatic Prompt Engineering: An Optimization Perspective
von: Li, Wenwu, et al.
Veröffentlicht: (2025)
von: Li, Wenwu, et al.
Veröffentlicht: (2025)
Structured Agent Distillation for Large Language Model
von: Liu, Jun, et al.
Veröffentlicht: (2025)
von: Liu, Jun, et al.
Veröffentlicht: (2025)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
von: Luo, Jun, et al.
Veröffentlicht: (2024)
von: Luo, Jun, et al.
Veröffentlicht: (2024)
Fine-Tuning is Subgraph Search: A New Lens on Learning Dynamics
von: Li, Yueyan, et al.
Veröffentlicht: (2025)
von: Li, Yueyan, et al.
Veröffentlicht: (2025)
Fighting Spurious Correlations in Text Classification via a Causal Learning Perspective
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
Unlocking the Potentials of Retrieval-Augmented Generation for Diffusion Language Models
von: Yu, Chuanyue, et al.
Veröffentlicht: (2026)
von: Yu, Chuanyue, et al.
Veröffentlicht: (2026)
Driving Intelligent IoT Monitoring and Control through Cloud Computing and Machine Learning
von: Li, Hanzhe, et al.
Veröffentlicht: (2024)
von: Li, Hanzhe, et al.
Veröffentlicht: (2024)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Where Paths Collide: A Comprehensive Survey of Classic and Learning-Based Multi-Agent Pathfinding
von: Wang, Shiyue, et al.
Veröffentlicht: (2025)
von: Wang, Shiyue, et al.
Veröffentlicht: (2025)
Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference
von: Song, Yuxuan, et al.
Veröffentlicht: (2025)
von: Song, Yuxuan, et al.
Veröffentlicht: (2025)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
von: Wang, Duo, et al.
Veröffentlicht: (2024)
von: Wang, Duo, et al.
Veröffentlicht: (2024)
Co-Evolution of Policy and Internal Reward for Language Agents
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
Continual Learning for Large Language Models: A Survey
von: Wu, Tongtong, et al.
Veröffentlicht: (2024)
von: Wu, Tongtong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scalable Reinforcement Learning for Virtual Machine Scheduling
von: Sheng, Junjie, et al.
Veröffentlicht: (2025) -
TextAtari: 100K Frames Game Playing with Language Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2025) -
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
von: Wu, Jiehao, et al.
Veröffentlicht: (2026) -
Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2026) -
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)