How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Yujian, Ji, Jiabao, An, Li, Jaakkola, Tommi, Zhang, Yang, Chang, Shiyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
por: Liu, Yujian, et al.
Publicado: (2025)
por: Liu, Yujian, et al.
Publicado: (2025)
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
por: Liu, Yujian, et al.
Publicado: (2024)
por: Liu, Yujian, et al.
Publicado: (2024)
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
por: Liu, Yujian, et al.
Publicado: (2024)
por: Liu, Yujian, et al.
Publicado: (2024)
Correcting Diffusion Generation through Resampling
por: Liu, Yujian, et al.
Publicado: (2023)
por: Liu, Yujian, et al.
Publicado: (2023)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
por: Ji, Jiabao, et al.
Publicado: (2024)
por: Ji, Jiabao, et al.
Publicado: (2024)
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
por: Su, Xuanbo, et al.
Publicado: (2026)
por: Su, Xuanbo, et al.
Publicado: (2026)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
por: Ji, Jiabao, et al.
Publicado: (2025)
por: Ji, Jiabao, et al.
Publicado: (2025)
ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
por: Hou, Bairu, et al.
Publicado: (2025)
por: Hou, Bairu, et al.
Publicado: (2025)
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
por: Ma, Ziyu, et al.
Publicado: (2026)
por: Ma, Ziyu, et al.
Publicado: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
por: An, Li, et al.
Publicado: (2025)
por: An, Li, et al.
Publicado: (2025)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
por: Tan, Rongyuan, et al.
Publicado: (2026)
por: Tan, Rongyuan, et al.
Publicado: (2026)
Augment before You Try: Knowledge-Enhanced Table Question Answering via Table Expansion
por: Liu, Yujian, et al.
Publicado: (2024)
por: Liu, Yujian, et al.
Publicado: (2024)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
por: Herandi, Amirhossein, et al.
Publicado: (2024)
por: Herandi, Amirhossein, et al.
Publicado: (2024)
Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents
por: Li, Xinze, et al.
Publicado: (2026)
por: Li, Xinze, et al.
Publicado: (2026)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
por: Che, Xinyu, et al.
Publicado: (2026)
por: Che, Xinyu, et al.
Publicado: (2026)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
por: Zhao, Yixiu, et al.
Publicado: (2025)
por: Zhao, Yixiu, et al.
Publicado: (2025)
Inducing Programmatic Skills for Agentic Tasks
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
Benchmarking LLM Tool-Use in the Wild
por: Yu, Peijie, et al.
Publicado: (2026)
por: Yu, Peijie, et al.
Publicado: (2026)
Improving Academic Skills Assessment with NLP and Ensemble Learning
por: Huang, Xinyi, et al.
Publicado: (2024)
por: Huang, Xinyi, et al.
Publicado: (2024)
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning
por: Zhu, Jiapeng, et al.
Publicado: (2026)
por: Zhu, Jiapeng, et al.
Publicado: (2026)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
por: Wang, Jingxing, et al.
Publicado: (2026)
por: Wang, Jingxing, et al.
Publicado: (2026)
SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
por: Pan, Shuai, et al.
Publicado: (2026)
por: Pan, Shuai, et al.
Publicado: (2026)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
por: Liu, Yi, et al.
Publicado: (2026)
por: Liu, Yi, et al.
Publicado: (2026)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
por: Chen, Shiqi, et al.
Publicado: (2026)
por: Chen, Shiqi, et al.
Publicado: (2026)
Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
por: Jiang, Pengcheng, et al.
Publicado: (2025)
por: Jiang, Pengcheng, et al.
Publicado: (2025)
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
por: Hou, Bairu, et al.
Publicado: (2023)
por: Hou, Bairu, et al.
Publicado: (2023)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
por: Li, Xiaoyuan, et al.
Publicado: (2026)
por: Li, Xiaoyuan, et al.
Publicado: (2026)
RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis
por: Tao, Zhengwei, et al.
Publicado: (2026)
por: Tao, Zhengwei, et al.
Publicado: (2026)
Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning
por: Shen, Junhao, et al.
Publicado: (2026)
por: Shen, Junhao, et al.
Publicado: (2026)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
por: Yu, Jianxiang, et al.
Publicado: (2026)
por: Yu, Jianxiang, et al.
Publicado: (2026)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
por: Chen, Chishui, et al.
Publicado: (2026)
por: Chen, Chishui, et al.
Publicado: (2026)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
por: Liu, Simiao, et al.
Publicado: (2026)
por: Liu, Simiao, et al.
Publicado: (2026)
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
por: Li, Hao, et al.
Publicado: (2026)
por: Li, Hao, et al.
Publicado: (2026)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
por: Nottingham, Kolby, et al.
Publicado: (2024)
por: Nottingham, Kolby, et al.
Publicado: (2024)
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
por: Liang, Qiliang, et al.
Publicado: (2026)
por: Liang, Qiliang, et al.
Publicado: (2026)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
por: Styles, Olly, et al.
Publicado: (2024)
por: Styles, Olly, et al.
Publicado: (2024)
Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub
por: Hu, Haichuan, et al.
Publicado: (2026)
por: Hu, Haichuan, et al.
Publicado: (2026)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
por: Wu, Minghao, et al.
Publicado: (2024)
por: Wu, Minghao, et al.
Publicado: (2024)
Job Skill Extraction via LLM-Centric Multi-Module Framework
por: Li, Guojing, et al.
Publicado: (2026)
por: Li, Guojing, et al.
Publicado: (2026)
Ejemplares similares
-
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
por: Liu, Yujian, et al.
Publicado: (2025) -
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
por: Liu, Yujian, et al.
Publicado: (2024) -
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
por: Liu, Yujian, et al.
Publicado: (2024) -
Correcting Diffusion Generation through Resampling
por: Liu, Yujian, et al.
Publicado: (2023) -
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
por: Ji, Jiabao, et al.
Publicado: (2024)