Saved in:
| Main Authors: | Sithakoul, Samuel, Meftah, Sara, Feutry, Clément |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.19897 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025)
by: Wang, Jianghui, et al.
Published: (2025)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024)
by: Böck, Adrian Jaques, et al.
Published: (2024)
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
by: Abdali, Sara, et al.
Published: (2024)
by: Abdali, Sara, et al.
Published: (2024)
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
by: Laurent, Jon M, et al.
Published: (2026)
by: Laurent, Jon M, et al.
Published: (2026)
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025)
by: Akimov, Farkhad, et al.
Published: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)
by: Stanovsky, Gabriel, et al.
Published: (2025)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
by: Azarafrooz, Ari
Published: (2026)
by: Azarafrooz, Ari
Published: (2026)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
by: Shin, Hagyeong, et al.
Published: (2025)
by: Shin, Hagyeong, et al.
Published: (2025)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
by: Mills, Edmund, et al.
Published: (2023)
by: Mills, Edmund, et al.
Published: (2023)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
by: Ji, Haonian, et al.
Published: (2026)
by: Ji, Haonian, et al.
Published: (2026)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
by: Arcadinho, Samuel, et al.
Published: (2024)
by: Arcadinho, Samuel, et al.
Published: (2024)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
by: Miyai, Atsuyuki, et al.
Published: (2026)
by: Miyai, Atsuyuki, et al.
Published: (2026)
A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI
by: Malinga, Melusi, et al.
Published: (2024)
by: Malinga, Melusi, et al.
Published: (2024)
A Quantum Inspired Variational Kernel and Explainable AI Framework for Cross Region Solar and Wind Energy Forecasting
by: Manjunath, Pavan, et al.
Published: (2026)
by: Manjunath, Pavan, et al.
Published: (2026)
Fine-Tuning and Evaluating Conversational AI for Agricultural Advisory
by: Singh, Sanyam, et al.
Published: (2026)
by: Singh, Sanyam, et al.
Published: (2026)
Explingo: Explaining AI Predictions using Large Language Models
by: Zytek, Alexandra, et al.
Published: (2024)
by: Zytek, Alexandra, et al.
Published: (2024)
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
by: Zhang, Jifan, et al.
Published: (2024)
by: Zhang, Jifan, et al.
Published: (2024)
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
by: Noever, David, et al.
Published: (2025)
by: Noever, David, et al.
Published: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
A Survey of the State of Explainable AI for Natural Language Processing
by: Danilevsky, Marina, et al.
Published: (2020)
by: Danilevsky, Marina, et al.
Published: (2020)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
by: Tice, Cameron, et al.
Published: (2026)
by: Tice, Cameron, et al.
Published: (2026)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
by: Sharma, Archit, et al.
Published: (2024)
by: Sharma, Archit, et al.
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
by: Karagoz, Atahan
Published: (2026)
by: Karagoz, Atahan
Published: (2026)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
AI and Generative AI for Research Discovery and Summarization
by: Glickman, Mark, et al.
Published: (2024)
by: Glickman, Mark, et al.
Published: (2024)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Towards Conversational Diagnostic AI
by: Tu, Tao, et al.
Published: (2024)
by: Tu, Tao, et al.
Published: (2024)
Towards Conversational AI for Disease Management
by: Palepu, Anil, et al.
Published: (2025)
by: Palepu, Anil, et al.
Published: (2025)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
by: Levi, Elad, et al.
Published: (2025)
by: Levi, Elad, et al.
Published: (2025)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
AI-rithmetic
by: Bie, Alex, et al.
Published: (2026)
by: Bie, Alex, et al.
Published: (2026)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Similar Items
-
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025) -
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025) -
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024) -
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
by: Abdali, Sara, et al.
Published: (2024) -
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
by: Laurent, Jon M, et al.
Published: (2026)