How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yujian, Ji, Jiabao, An, Li, Jaakkola, Tommi, Zhang, Yang, Chang, Shiyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
di: Liu, Yujian, et al.
Pubblicazione: (2025)
di: Liu, Yujian, et al.
Pubblicazione: (2025)
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
di: Liu, Yujian, et al.
Pubblicazione: (2024)
di: Liu, Yujian, et al.
Pubblicazione: (2024)
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
di: Liu, Yujian, et al.
Pubblicazione: (2024)
di: Liu, Yujian, et al.
Pubblicazione: (2024)
Correcting Diffusion Generation through Resampling
di: Liu, Yujian, et al.
Pubblicazione: (2023)
di: Liu, Yujian, et al.
Pubblicazione: (2023)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
di: Ji, Jiabao, et al.
Pubblicazione: (2025)
di: Ji, Jiabao, et al.
Pubblicazione: (2025)
ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
di: Hou, Bairu, et al.
Pubblicazione: (2025)
di: Hou, Bairu, et al.
Pubblicazione: (2025)
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
di: Ma, Ziyu, et al.
Pubblicazione: (2026)
di: Ma, Ziyu, et al.
Pubblicazione: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
di: An, Li, et al.
Pubblicazione: (2025)
di: An, Li, et al.
Pubblicazione: (2025)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
Augment before You Try: Knowledge-Enhanced Table Question Answering via Table Expansion
di: Liu, Yujian, et al.
Pubblicazione: (2024)
di: Liu, Yujian, et al.
Pubblicazione: (2024)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
di: Herandi, Amirhossein, et al.
Pubblicazione: (2024)
di: Herandi, Amirhossein, et al.
Pubblicazione: (2024)
Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents
di: Li, Xinze, et al.
Pubblicazione: (2026)
di: Li, Xinze, et al.
Pubblicazione: (2026)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
di: Che, Xinyu, et al.
Pubblicazione: (2026)
di: Che, Xinyu, et al.
Pubblicazione: (2026)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
di: Zhao, Yixiu, et al.
Pubblicazione: (2025)
di: Zhao, Yixiu, et al.
Pubblicazione: (2025)
Inducing Programmatic Skills for Agentic Tasks
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
Benchmarking LLM Tool-Use in the Wild
di: Yu, Peijie, et al.
Pubblicazione: (2026)
di: Yu, Peijie, et al.
Pubblicazione: (2026)
Improving Academic Skills Assessment with NLP and Ensemble Learning
di: Huang, Xinyi, et al.
Pubblicazione: (2024)
di: Huang, Xinyi, et al.
Pubblicazione: (2024)
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning
di: Zhu, Jiapeng, et al.
Pubblicazione: (2026)
di: Zhu, Jiapeng, et al.
Pubblicazione: (2026)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
di: Wang, Jingxing, et al.
Pubblicazione: (2026)
di: Wang, Jingxing, et al.
Pubblicazione: (2026)
SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
di: Pan, Shuai, et al.
Pubblicazione: (2026)
di: Pan, Shuai, et al.
Pubblicazione: (2026)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
di: Liu, Yi, et al.
Pubblicazione: (2026)
di: Liu, Yi, et al.
Pubblicazione: (2026)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
di: Hou, Bairu, et al.
Pubblicazione: (2023)
di: Hou, Bairu, et al.
Pubblicazione: (2023)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis
di: Tao, Zhengwei, et al.
Pubblicazione: (2026)
di: Tao, Zhengwei, et al.
Pubblicazione: (2026)
Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning
di: Shen, Junhao, et al.
Pubblicazione: (2026)
di: Shen, Junhao, et al.
Pubblicazione: (2026)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
di: Yu, Jianxiang, et al.
Pubblicazione: (2026)
di: Yu, Jianxiang, et al.
Pubblicazione: (2026)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
di: Chen, Chishui, et al.
Pubblicazione: (2026)
di: Chen, Chishui, et al.
Pubblicazione: (2026)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
di: Liu, Simiao, et al.
Pubblicazione: (2026)
di: Liu, Simiao, et al.
Pubblicazione: (2026)
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
di: Nottingham, Kolby, et al.
Pubblicazione: (2024)
di: Nottingham, Kolby, et al.
Pubblicazione: (2024)
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
di: Liang, Qiliang, et al.
Pubblicazione: (2026)
di: Liang, Qiliang, et al.
Pubblicazione: (2026)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
di: Styles, Olly, et al.
Pubblicazione: (2024)
di: Styles, Olly, et al.
Pubblicazione: (2024)
Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub
di: Hu, Haichuan, et al.
Pubblicazione: (2026)
di: Hu, Haichuan, et al.
Pubblicazione: (2026)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
di: Wu, Minghao, et al.
Pubblicazione: (2024)
di: Wu, Minghao, et al.
Pubblicazione: (2024)
Job Skill Extraction via LLM-Centric Multi-Module Framework
di: Li, Guojing, et al.
Pubblicazione: (2026)
di: Li, Guojing, et al.
Pubblicazione: (2026)
Documenti analoghi
-
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
di: Liu, Yujian, et al.
Pubblicazione: (2025) -
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
di: Liu, Yujian, et al.
Pubblicazione: (2024) -
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
di: Liu, Yujian, et al.
Pubblicazione: (2024) -
Correcting Diffusion Generation through Resampling
di: Liu, Yujian, et al.
Pubblicazione: (2023) -
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
di: Ji, Jiabao, et al.
Pubblicazione: (2024)