Saved in:
| Main Authors: | Provost, Marc-Antoine, Ilenic, Nejc, Solinas, Christopher, Beardsell, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.23660 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformer Based Planning in the Observation Space with Applications to Trick Taking Card Games
by: Rebstock, Douglas, et al.
Published: (2024)
by: Rebstock, Douglas, et al.
Published: (2024)
SEER: The Span-based Emotion Evidence Retrieval Benchmark
by: Sampath, Aneesha, et al.
Published: (2025)
by: Sampath, Aneesha, et al.
Published: (2025)
On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
by: Fang, Jingchao, et al.
Published: (2024)
by: Fang, Jingchao, et al.
Published: (2024)
Adaptive Wizard for Removing Cross-Tier Misconfigurations in Active Directory
by: Ngo, Huy Q., et al.
Published: (2025)
by: Ngo, Huy Q., et al.
Published: (2025)
PromptWizard: Task-Aware Prompt Optimization Framework
by: Agarwal, Eshaan, et al.
Published: (2024)
by: Agarwal, Eshaan, et al.
Published: (2024)
WizardCoder: Empowering Code Large Language Models with Evol-Instruct
by: Luo, Ziyang, et al.
Published: (2023)
by: Luo, Ziyang, et al.
Published: (2023)
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
by: Ferreira, Rafael, et al.
Published: (2023)
by: Ferreira, Rafael, et al.
Published: (2023)
The Carbon Footprint Wizard: A Knowledge-Augmented AI Interface for Streamlining Food Carbon Footprint Analysis
by: Aslan, Mustafa Kaan, et al.
Published: (2025)
by: Aslan, Mustafa Kaan, et al.
Published: (2025)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
RecWizard: A Toolkit for Conversational Recommendation with Modular, Portable Models and Interactive User Interface
by: Zhang, Zeyuan, et al.
Published: (2024)
by: Zhang, Zeyuan, et al.
Published: (2024)
What You Feel Is Not What They See: On Predicting Self-Reported Emotion from Third-Party Observer Labels
by: El-Tawil, Yara, et al.
Published: (2026)
by: El-Tawil, Yara, et al.
Published: (2026)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
by: Luo, Haipeng, et al.
Published: (2023)
by: Luo, Haipeng, et al.
Published: (2023)
Neural Bayesian Filtering
by: Solinas, Christopher, et al.
Published: (2025)
by: Solinas, Christopher, et al.
Published: (2025)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
by: Sampath, Aneesha, et al.
Published: (2025)
by: Sampath, Aneesha, et al.
Published: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
Learning-based Multi-agent Race Strategies in Formula 1
by: Fieni, Giona, et al.
Published: (2026)
by: Fieni, Giona, et al.
Published: (2026)
Gaussian Match-and-Copy: A Minimalist Benchmark for Studying Transformer Induction
by: Gonon, Antoine, et al.
Published: (2026)
by: Gonon, Antoine, et al.
Published: (2026)
An Extended Jump Functions Benchmark for the Analysis of Randomized Search Heuristics
by: Bambury, Henry, et al.
Published: (2021)
by: Bambury, Henry, et al.
Published: (2021)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
by: Allard, Marc-Antoine, et al.
Published: (2024)
by: Allard, Marc-Antoine, et al.
Published: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Centralized vs. Decentralized Security for Space AI Systems? A New Look
by: Schmitt, Noam, et al.
Published: (2025)
by: Schmitt, Noam, et al.
Published: (2025)
Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings
by: Lagesse, Adrien, et al.
Published: (2025)
by: Lagesse, Adrien, et al.
Published: (2025)
Experiential Reflective Learning for Self-Improving LLM Agents
by: Allard, Marc-Antoine, et al.
Published: (2026)
by: Allard, Marc-Antoine, et al.
Published: (2026)
SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities
by: Lalou, Yanis, et al.
Published: (2024)
by: Lalou, Yanis, et al.
Published: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
Benchmarking Agents in Insurance Underwriting Environments
by: Dsouza, Amanda, et al.
Published: (2026)
by: Dsouza, Amanda, et al.
Published: (2026)
German Text Embedding Clustering Benchmark
by: Wehrli, Silvan, et al.
Published: (2024)
by: Wehrli, Silvan, et al.
Published: (2024)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
by: Stoisser, Josefa Lia, et al.
Published: (2026)
by: Stoisser, Josefa Lia, et al.
Published: (2026)
Benchmarking Domain Adaptation for Chemical Processes on the Tennessee Eastman Process
by: Montesuma, Eduardo Fernandes, et al.
Published: (2023)
by: Montesuma, Eduardo Fernandes, et al.
Published: (2023)
Revisiting Synthetic Human Trajectories: Imitative Generation and Benchmarks Beyond Datasaurus
by: Deng, Bangchao, et al.
Published: (2024)
by: Deng, Bangchao, et al.
Published: (2024)
A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
by: Morio, Gaku, et al.
Published: (2025)
by: Morio, Gaku, et al.
Published: (2025)
Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks
by: Bechler-Speicher, Maya, et al.
Published: (2025)
by: Bechler-Speicher, Maya, et al.
Published: (2025)
Towards Accurate Forecasting of Renewable Energy : Building Datasets and Benchmarking Machine Learning Models for Solar and Wind Power in France
by: Lindas, Eloi, et al.
Published: (2025)
by: Lindas, Eloi, et al.
Published: (2025)
A network analysis of decision strategies of human experts in steel manufacturing
by: Merten, Daniel Christopher, et al.
Published: (2021)
by: Merten, Daniel Christopher, et al.
Published: (2021)
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations
by: Marchesini, Enrico, et al.
Published: (2025)
by: Marchesini, Enrico, et al.
Published: (2025)
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
by: Zanoli, Christopher, et al.
Published: (2026)
by: Zanoli, Christopher, et al.
Published: (2026)
DStruct2Design: Data and Benchmarks for Data Structure Driven Generative Floor Plan Design
by: Luo, Zhi Hao, et al.
Published: (2024)
by: Luo, Zhi Hao, et al.
Published: (2024)
Similar Items
-
Transformer Based Planning in the Observation Space with Applications to Trick Taking Card Games
by: Rebstock, Douglas, et al.
Published: (2024) -
SEER: The Span-based Emotion Evidence Retrieval Benchmark
by: Sampath, Aneesha, et al.
Published: (2025) -
On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
by: Fang, Jingchao, et al.
Published: (2024) -
Adaptive Wizard for Removing Cross-Tier Misconfigurations in Active Directory
by: Ngo, Huy Q., et al.
Published: (2025) -
PromptWizard: Task-Aware Prompt Optimization Framework
by: Agarwal, Eshaan, et al.
Published: (2024)