Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants' API Invocation Capabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Mu, Honglin, Xu, Yang, Feng, Yunlong, Han, Xiaofeng, Li, Yitong, Hou, Yutai, Che, Wanxiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Concise and Precise Context Compression for Tool-Using Language Models
por: Xu, Yang, et al.
Publicado: (2024)
por: Xu, Yang, et al.
Publicado: (2024)
Language Anisotropic Cross-Lingual Model Editing
por: Xu, Yang, et al.
Publicado: (2022)
por: Xu, Yang, et al.
Publicado: (2022)
Self-Constructed Context Decompilation with Fined-grained Alignment Enhancement
por: Feng, Yunlong, et al.
Publicado: (2024)
por: Feng, Yunlong, et al.
Publicado: (2024)
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
por: Mu, Honglin, et al.
Publicado: (2025)
por: Mu, Honglin, et al.
Publicado: (2025)
Improving Language Model Reasoning with Self-motivated Learning
por: Feng, Yunlong, et al.
Publicado: (2024)
por: Feng, Yunlong, et al.
Publicado: (2024)
Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
por: Mu, Honglin, et al.
Publicado: (2024)
por: Mu, Honglin, et al.
Publicado: (2024)
A Two-Stage Framework with Self-Supervised Distillation For Cross-Domain Text Classification
por: Feng, Yunlong, et al.
Publicado: (2023)
por: Feng, Yunlong, et al.
Publicado: (2023)
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
por: Hou, Nan, et al.
Publicado: (2026)
por: Hou, Nan, et al.
Publicado: (2026)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
por: Wang, Dingzriui, et al.
Publicado: (2025)
por: Wang, Dingzriui, et al.
Publicado: (2025)
Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
por: Wang, Dingzirui, et al.
Publicado: (2025)
por: Wang, Dingzirui, et al.
Publicado: (2025)
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
por: Chen, Qiguang, et al.
Publicado: (2024)
por: Chen, Qiguang, et al.
Publicado: (2024)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
por: Xu, Yuzhuang, et al.
Publicado: (2026)
por: Xu, Yuzhuang, et al.
Publicado: (2026)
HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
por: Shan, Baocai, et al.
Publicado: (2026)
por: Shan, Baocai, et al.
Publicado: (2026)
When Does Context Help? Error Dynamics of Contextual Information in Large Language Models
por: Wang, Dingzirui, et al.
Publicado: (2026)
por: Wang, Dingzirui, et al.
Publicado: (2026)
Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification
por: Li, Bohan, et al.
Publicado: (2023)
por: Li, Bohan, et al.
Publicado: (2023)
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
por: Xu, Yuzhuang, et al.
Publicado: (2026)
por: Xu, Yuzhuang, et al.
Publicado: (2026)
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
por: Wang, Renxi, et al.
Publicado: (2025)
por: Wang, Renxi, et al.
Publicado: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
por: Shi, Zhichao, et al.
Publicado: (2025)
por: Shi, Zhichao, et al.
Publicado: (2025)
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
por: Wang, Yixuan, et al.
Publicado: (2024)
por: Wang, Yixuan, et al.
Publicado: (2024)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
por: Lin, Lizhi, et al.
Publicado: (2024)
por: Lin, Lizhi, et al.
Publicado: (2024)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
por: Huang, Xu, et al.
Publicado: (2025)
por: Huang, Xu, et al.
Publicado: (2025)
Beyond Browsing: API-Based Web Agents
por: Song, Yueqi, et al.
Publicado: (2024)
por: Song, Yueqi, et al.
Publicado: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2024)
por: Kim, Eunsu, et al.
Publicado: (2024)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
por: Xu, Yuzhuang, et al.
Publicado: (2024)
por: Xu, Yuzhuang, et al.
Publicado: (2024)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
por: Niu, Tianhao, et al.
Publicado: (2026)
por: Niu, Tianhao, et al.
Publicado: (2026)
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation
por: Dai, Jianbo, et al.
Publicado: (2024)
por: Dai, Jianbo, et al.
Publicado: (2024)
Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL
por: Wang, Yihan, et al.
Publicado: (2026)
por: Wang, Yihan, et al.
Publicado: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
por: Kim, Doyoung, et al.
Publicado: (2026)
por: Kim, Doyoung, et al.
Publicado: (2026)
OneBit: Towards Extremely Low-bit Large Language Models
por: Xu, Yuzhuang, et al.
Publicado: (2024)
por: Xu, Yuzhuang, et al.
Publicado: (2024)
Scaling Laws for Agent Harnesses via Effective Feedback Compute
por: Zhang, Xuanliang, et al.
Publicado: (2026)
por: Zhang, Xuanliang, et al.
Publicado: (2026)
How Do Language Models Understand Tables? A Mechanistic Analysis of Cell Location
por: Zhang, Xuanliang, et al.
Publicado: (2026)
por: Zhang, Xuanliang, et al.
Publicado: (2026)
Abacus-SQL: A Text-to-SQL System Empowering Cross-Domain and Open-Domain Database Retrieval
por: Xu, Keyan, et al.
Publicado: (2025)
por: Xu, Keyan, et al.
Publicado: (2025)
Pro-HAN: A Heterogeneous Graph Attention Network for Profile-Based Spoken Language Understanding
por: Teng, Dechuan, et al.
Publicado: (2024)
por: Teng, Dechuan, et al.
Publicado: (2024)
RoT: Enhancing Table Reasoning with Iterative Row-Wise Traversals
por: Zhang, Xuanliang, et al.
Publicado: (2025)
por: Zhang, Xuanliang, et al.
Publicado: (2025)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
por: Zhang, Xuanliang, et al.
Publicado: (2025)
por: Zhang, Xuanliang, et al.
Publicado: (2025)
AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution
por: Wang, Dingzirui, et al.
Publicado: (2025)
por: Wang, Dingzirui, et al.
Publicado: (2025)
How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix Perspective
por: Luo, Xianzhen, et al.
Publicado: (2025)
por: Luo, Xianzhen, et al.
Publicado: (2025)
Semi-Instruct: Bridging Natural-Instruct and Self-Instruct for Code Large Language Models
por: Luo, Xianzhen, et al.
Publicado: (2024)
por: Luo, Xianzhen, et al.
Publicado: (2024)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
por: Chen, Qiguang, et al.
Publicado: (2025)
por: Chen, Qiguang, et al.
Publicado: (2025)
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
por: Wang, Dingzirui, et al.
Publicado: (2025)
por: Wang, Dingzirui, et al.
Publicado: (2025)
Ejemplares similares
-
Concise and Precise Context Compression for Tool-Using Language Models
por: Xu, Yang, et al.
Publicado: (2024) -
Language Anisotropic Cross-Lingual Model Editing
por: Xu, Yang, et al.
Publicado: (2022) -
Self-Constructed Context Decompilation with Fined-grained Alignment Enhancement
por: Feng, Yunlong, et al.
Publicado: (2024) -
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
por: Mu, Honglin, et al.
Publicado: (2025) -
Improving Language Model Reasoning with Self-motivated Learning
por: Feng, Yunlong, et al.
Publicado: (2024)