Gespeichert in:
| Hauptverfasser: | Li, Yanming, Zhang, Xuelin, Lu, WenJie, Tang, Ziye, Wu, Maodong, Luo, Haotian, Wu, Tongtong, Peng, Zijie, Mi, Hongze, Feng, Yibo, Tan, Naiqiang, Huang, Chao, Chen, Hong, Shen, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.08335 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution
von: Mi, Hongze, et al.
Veröffentlicht: (2026)
von: Mi, Hongze, et al.
Veröffentlicht: (2026)
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
von: Mi, Hongze, et al.
Veröffentlicht: (2025)
von: Mi, Hongze, et al.
Veröffentlicht: (2025)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
Who Deserves the Credit for Lower Unemployment? Structural Monetary Policy Tools and Corporate Labour Employment in China
von: Xue Li, et al.
Veröffentlicht: (2025)
von: Xue Li, et al.
Veröffentlicht: (2025)
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
Comparative Outcomes of Different Surgical Approaches for Non‐Lactational Mastitis With Posterior and Non‐Posterior Space Abscesses: A Retrospective Cohort Study
von: WenJie Zhang, et al.
Veröffentlicht: (2025)
von: WenJie Zhang, et al.
Veröffentlicht: (2025)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
CARD: Towards Conditional Design of Multi-agent Topological Structures
von: Wu, Tongtong, et al.
Veröffentlicht: (2026)
von: Wu, Tongtong, et al.
Veröffentlicht: (2026)
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents
von: Hua, Yun, et al.
Veröffentlicht: (2025)
von: Hua, Yun, et al.
Veröffentlicht: (2025)
What Deserves Memory: Adaptive Memory Distillation for LLM Agents
von: Ma, Wenquan, et al.
Veröffentlicht: (2025)
von: Ma, Wenquan, et al.
Veröffentlicht: (2025)
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Bag of Tricks for Inference-time Computation of LLM Reasoning
von: Liu, Fan, et al.
Veröffentlicht: (2025)
von: Liu, Fan, et al.
Veröffentlicht: (2025)
Improving DAPO from a Mixed-Policy Perspective
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
von: He, Jessica, et al.
Veröffentlicht: (2025)
von: He, Jessica, et al.
Veröffentlicht: (2025)
FusionRegister: Every Infrared and Visible Image Fusion Deserves Registration
von: Bian, Congcong, et al.
Veröffentlicht: (2026)
von: Bian, Congcong, et al.
Veröffentlicht: (2026)
Prompt-Driven Low-Altitude Edge Intelligence: Modular Agents and Generative Reasoning
von: You, Jiahao, et al.
Veröffentlicht: (2026)
von: You, Jiahao, et al.
Veröffentlicht: (2026)
Enhancing Target-Guided Proactive Dialogue Systems via Conversational Scenario Modeling and Intent-Keyword Bridging
von: Li, Maodong, et al.
Veröffentlicht: (2026)
von: Li, Maodong, et al.
Veröffentlicht: (2026)
Who Deserves to Stay? Latent Profiles of Public Perceptions of Migrant Deservingness in Turkey
von: Dilara Turgut, et al.
Veröffentlicht: (2025)
von: Dilara Turgut, et al.
Veröffentlicht: (2025)
Multi-Hop Question Generation via Dual-Perspective Keyword Guidance
von: Li, Maodong, et al.
Veröffentlicht: (2025)
von: Li, Maodong, et al.
Veröffentlicht: (2025)
SCAR: Shapley Credit Assignment for More Efficient RLHF
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Hindsight Credit Assignment for Long-Horizon LLM Agents
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
von: Wang, Bin, et al.
Veröffentlicht: (2025)
von: Wang, Bin, et al.
Veröffentlicht: (2025)
SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
A Historical Interaction-Enhanced Shapley Policy Gradient Algorithm for Multi-Agent Credit Assignment
von: Ding, Ao, et al.
Veröffentlicht: (2025)
von: Ding, Ao, et al.
Veröffentlicht: (2025)
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
von: Cheng, Jie, et al.
Veröffentlicht: (2025)
von: Cheng, Jie, et al.
Veröffentlicht: (2025)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
SHARP: Experiences in Library Automation
von: Smith, Ruth Camp
Veröffentlicht: (1974)
von: Smith, Ruth Camp
Veröffentlicht: (1974)
Deserving the Option to Give
von: Jeffrey Moriarty
Veröffentlicht: (2024)
von: Jeffrey Moriarty
Veröffentlicht: (2024)
Random ISAC Signals Deserve Dedicated Precoding
von: Lu, Shihang, et al.
Veröffentlicht: (2023)
von: Lu, Shihang, et al.
Veröffentlicht: (2023)
Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate
von: Wu, Chenhao, et al.
Veröffentlicht: (2025)
von: Wu, Chenhao, et al.
Veröffentlicht: (2025)
Who Deserves Scarce Health and Education Resources? How Policy Context Shapes Target Group Deservingness
von: Elizabeth Bell, et al.
Veröffentlicht: (2025)
von: Elizabeth Bell, et al.
Veröffentlicht: (2025)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues
von: Kang, Xinyue, et al.
Veröffentlicht: (2026)
von: Kang, Xinyue, et al.
Veröffentlicht: (2026)
After Returning to the Rural: The (Un)Sustainable Reintegration of Internal Migrant Workers in China
von: Mengyao Cheng, et al.
Veröffentlicht: (2025)
von: Mengyao Cheng, et al.
Veröffentlicht: (2025)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution
von: Mi, Hongze, et al.
Veröffentlicht: (2026) -
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
von: Mi, Hongze, et al.
Veröffentlicht: (2025) -
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
von: Luo, Haotian, et al.
Veröffentlicht: (2025) -
Who Deserves the Credit for Lower Unemployment? Structural Monetary Policy Tools and Corporate Labour Employment in China
von: Xue Li, et al.
Veröffentlicht: (2025) -
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
von: Luo, Haotian, et al.
Veröffentlicht: (2025)