Don't Just Fine-tune the Agent, Tune the Environment
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Siyuan, Wang, Zechuan, Zhang, Hongxuan, Wu, Qintong, Gan, Leilei, Zhuang, Chenyi, Gu, Jinjie, Lin, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
di: Tan, Zhiwen, et al.
Pubblicazione: (2025)
di: Tan, Zhiwen, et al.
Pubblicazione: (2025)
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
di: Xie, Zhitian, et al.
Pubblicazione: (2025)
di: Xie, Zhitian, et al.
Pubblicazione: (2025)
SQLCritic: Correcting Text-to-SQL Generation via Clause-wise Critic
di: Chen, Jikai, et al.
Pubblicazione: (2025)
di: Chen, Jikai, et al.
Pubblicazione: (2025)
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
di: He, Kaiwen, et al.
Pubblicazione: (2025)
di: He, Kaiwen, et al.
Pubblicazione: (2025)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
di: Chen, Jikai, et al.
Pubblicazione: (2026)
di: Chen, Jikai, et al.
Pubblicazione: (2026)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
di: Chen, Jikai, et al.
Pubblicazione: (2025)
di: Chen, Jikai, et al.
Pubblicazione: (2025)
AWorld: Orchestrating the Training Recipe for Agentic AI
di: Yu, Chengyue, et al.
Pubblicazione: (2025)
di: Yu, Chengyue, et al.
Pubblicazione: (2025)
LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
di: Zhang, Kexun, et al.
Pubblicazione: (2023)
di: Zhang, Kexun, et al.
Pubblicazione: (2023)
FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
di: Xu, Zengzhuang, et al.
Pubblicazione: (2025)
di: Xu, Zengzhuang, et al.
Pubblicazione: (2025)
Large Pre-Training Datasets Don't Always Guarantee Robustness after Fine-Tuning
di: Hwang, Jaedong, et al.
Pubblicazione: (2024)
di: Hwang, Jaedong, et al.
Pubblicazione: (2024)
H2Tune: Federated Foundation Model Fine-Tuning with Hybrid Heterogeneity
di: Guo, Wei, et al.
Pubblicazione: (2025)
di: Guo, Wei, et al.
Pubblicazione: (2025)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM
di: Yu, Chengyue, et al.
Pubblicazione: (2024)
di: Yu, Chengyue, et al.
Pubblicazione: (2024)
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
di: Zhao, Yao, et al.
Pubblicazione: (2023)
di: Zhao, Yao, et al.
Pubblicazione: (2023)
Don't Just Demo, Teach Me the Principles: A Principle-Based Multi-Agent Prompting Strategy for Text Classification
di: Wei, Peipei, et al.
Pubblicazione: (2025)
di: Wei, Peipei, et al.
Pubblicazione: (2025)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
di: Xie, Zhitian, et al.
Pubblicazione: (2024)
di: Xie, Zhitian, et al.
Pubblicazione: (2024)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
di: Yu, Zhuoran, et al.
Pubblicazione: (2023)
di: Yu, Zhuoran, et al.
Pubblicazione: (2023)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2026)
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2026)
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning
di: He, Yongquan, et al.
Pubblicazione: (2024)
di: He, Yongquan, et al.
Pubblicazione: (2024)
Don't Just Translate, Agitate: Using Large Language Models as Devil's Advocates for AI Explanations
di: Suh, Ashley, et al.
Pubblicazione: (2025)
di: Suh, Ashley, et al.
Pubblicazione: (2025)
Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models
di: Zeng, Qingyuan, et al.
Pubblicazione: (2026)
di: Zeng, Qingyuan, et al.
Pubblicazione: (2026)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
di: Zhang, Jiefu, et al.
Pubblicazione: (2026)
di: Zhang, Jiefu, et al.
Pubblicazione: (2026)
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
di: Bian, Yutong, et al.
Pubblicazione: (2025)
di: Bian, Yutong, et al.
Pubblicazione: (2025)
Please Don't Kill My Vibe: Empowering Agents with Data Flow Control
di: Summers, Charlie, et al.
Pubblicazione: (2025)
di: Summers, Charlie, et al.
Pubblicazione: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
di: Lee, Joosung, et al.
Pubblicazione: (2026)
di: Lee, Joosung, et al.
Pubblicazione: (2026)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
di: Wang, Maolin, et al.
Pubblicazione: (2023)
di: Wang, Maolin, et al.
Pubblicazione: (2023)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
di: Otten, Daniel, et al.
Pubblicazione: (2026)
di: Otten, Daniel, et al.
Pubblicazione: (2026)
MorphAgent: Empowering Agents through Self-Evolving Profiles and Decentralized Collaboration
di: Lu, Siyuan, et al.
Pubblicazione: (2024)
di: Lu, Siyuan, et al.
Pubblicazione: (2024)
Formalize, Don't Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
s3: You Don't Need That Much Data to Train a Search Agent via RL
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
Forget the Data and Fine-Tuning! Just Fold the Network to Compress
di: Wang, Dong, et al.
Pubblicazione: (2025)
di: Wang, Dong, et al.
Pubblicazione: (2025)
DNS-Rec: Data-aware Neural Architecture Search for Recommender Systems
di: Zhang, Sheng, et al.
Pubblicazione: (2024)
di: Zhang, Sheng, et al.
Pubblicazione: (2024)
Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing
di: Liu, Wenhao, et al.
Pubblicazione: (2024)
di: Liu, Wenhao, et al.
Pubblicazione: (2024)
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2025)
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
Don't Get Too Excited -- Eliciting Emotions in LLMs
di: Fazzi, Gino Franco, et al.
Pubblicazione: (2025)
di: Fazzi, Gino Franco, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
di: Tan, Zhiwen, et al.
Pubblicazione: (2025) -
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
di: Xie, Zhitian, et al.
Pubblicazione: (2025) -
SQLCritic: Correcting Text-to-SQL Generation via Clause-wise Critic
di: Chen, Jikai, et al.
Pubblicazione: (2025) -
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
di: He, Kaiwen, et al.
Pubblicazione: (2025) -
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
di: Chen, Jikai, et al.
Pubblicazione: (2026)