Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Arcadinho, Samuel, Aparicio, David, Almeida, Mariana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
di: Singh, Chandan, et al.
Pubblicazione: (2026)
di: Singh, Chandan, et al.
Pubblicazione: (2026)
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
di: de Mijolla, Damien, et al.
Pubblicazione: (2024)
di: de Mijolla, Damien, et al.
Pubblicazione: (2024)
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
tinyBenchmarks: evaluating LLMs with fewer examples
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
di: Lu, Jun, et al.
Pubblicazione: (2024)
di: Lu, Jun, et al.
Pubblicazione: (2024)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
di: Lu, Chris, et al.
Pubblicazione: (2024)
di: Lu, Chris, et al.
Pubblicazione: (2024)
Towards physician-centered oversight of conversational diagnostic AI
di: Vedadi, Elahe, et al.
Pubblicazione: (2025)
di: Vedadi, Elahe, et al.
Pubblicazione: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
di: Lei, Ge, et al.
Pubblicazione: (2025)
di: Lei, Ge, et al.
Pubblicazione: (2025)
Towards Execution-Grounded Automated AI Research
di: Si, Chenglei, et al.
Pubblicazione: (2026)
di: Si, Chenglei, et al.
Pubblicazione: (2026)
Retrieval-augmented GUI Agents with Generative Guidelines
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
BEExAI: Benchmark to Evaluate Explainable AI
di: Sithakoul, Samuel, et al.
Pubblicazione: (2024)
di: Sithakoul, Samuel, et al.
Pubblicazione: (2024)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
AI Knowledge Assist: An Automated Approach for the Creation of Knowledge Bases for Conversational AI Agents
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
di: Maar, Jim, et al.
Pubblicazione: (2026)
di: Maar, Jim, et al.
Pubblicazione: (2026)
Automated Text Scoring in the Age of Generative AI for the GPU-poor
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
AIGS: Generating Science from AI-Powered Automated Falsification
di: Liu, Zijun, et al.
Pubblicazione: (2024)
di: Liu, Zijun, et al.
Pubblicazione: (2024)
Remote Labor Index: Measuring AI Automation of Remote Work
di: Mazeika, Mantas, et al.
Pubblicazione: (2025)
di: Mazeika, Mantas, et al.
Pubblicazione: (2025)
TxGemma: Efficient and Agentic LLMs for Therapeutics
di: Wang, Eric, et al.
Pubblicazione: (2025)
di: Wang, Eric, et al.
Pubblicazione: (2025)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
di: Yamada, Yutaro, et al.
Pubblicazione: (2025)
di: Yamada, Yutaro, et al.
Pubblicazione: (2025)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
di: Kachwala, Zoher, et al.
Pubblicazione: (2025)
di: Kachwala, Zoher, et al.
Pubblicazione: (2025)
Representation of perceived prosodic similarity of conversational feedback
di: Qian, Livia, et al.
Pubblicazione: (2025)
di: Qian, Livia, et al.
Pubblicazione: (2025)
TLDR at SemEval-2024 Task 2: T5-generated clinical-Language summaries for DeBERTa Report Analysis
di: Das, Spandan, et al.
Pubblicazione: (2024)
di: Das, Spandan, et al.
Pubblicazione: (2024)
RuAG: Learned-rule-augmented Generation for Large Language Models
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
di: Arasteh, Soroosh Tayebi, et al.
Pubblicazione: (2024)
di: Arasteh, Soroosh Tayebi, et al.
Pubblicazione: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
The AI Data Scientist
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
Aviary: training language agents on challenging scientific tasks
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
di: Tice, Cameron, et al.
Pubblicazione: (2026)
di: Tice, Cameron, et al.
Pubblicazione: (2026)
Speaking the Same Language: Leveraging LLMs in Standardizing Clinical Data for AI
di: Sett, Arindam, et al.
Pubblicazione: (2024)
di: Sett, Arindam, et al.
Pubblicazione: (2024)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
di: Casademunt, Helena, et al.
Pubblicazione: (2026)
di: Casademunt, Helena, et al.
Pubblicazione: (2026)
DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning
di: Guo, Xun, et al.
Pubblicazione: (2024)
di: Guo, Xun, et al.
Pubblicazione: (2024)
PersonaGym: Evaluating Persona Agents and LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
di: Brodeur, Peter, et al.
Pubblicazione: (2026)
di: Brodeur, Peter, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
di: Singh, Chandan, et al.
Pubblicazione: (2026) -
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
di: de Mijolla, Damien, et al.
Pubblicazione: (2024) -
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024) -
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025) -
tinyBenchmarks: evaluating LLMs with fewer examples
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)