Guardado en:
| Autores principales: | Wang, Paul, Garcia-Bourrée, Jade, Kermarrec, Anne-Marie, Corruble, Vincent |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.19377 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fairness Auditing with Multi-Agent Collaboration
por: de Vos, Martijn, et al.
Publicado: (2024)
por: de Vos, Martijn, et al.
Publicado: (2024)
Optimizing Agentic Workflows using Meta-tools
por: Abuzakuk, Sami, et al.
Publicado: (2026)
por: Abuzakuk, Sami, et al.
Publicado: (2026)
Robust ML Auditing using Prior Knowledge
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
Effective LoRA Adapter Routing using Task Representations
por: Dhasade, Akash, et al.
Publicado: (2026)
por: Dhasade, Akash, et al.
Publicado: (2026)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
por: Dorn, Diego, et al.
Publicado: (2024)
por: Dorn, Diego, et al.
Publicado: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
por: Yoa, Seungdong, et al.
Publicado: (2026)
por: Yoa, Seungdong, et al.
Publicado: (2026)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
por: Singh, Prabhant, et al.
Publicado: (2025)
por: Singh, Prabhant, et al.
Publicado: (2025)
[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
por: Pollo, Jorge Carrasco, et al.
Publicado: (2026)
por: Pollo, Jorge Carrasco, et al.
Publicado: (2026)
QuickDrop: Efficient Federated Unlearning by Integrated Dataset Distillation
por: Dhasade, Akash, et al.
Publicado: (2023)
por: Dhasade, Akash, et al.
Publicado: (2023)
Fast In-Spectrum Graph Watermarks
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
por: Liu, Hugh Xuechen, et al.
Publicado: (2026)
por: Liu, Hugh Xuechen, et al.
Publicado: (2026)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
por: Yuan, Peiwen, et al.
Publicado: (2025)
por: Yuan, Peiwen, et al.
Publicado: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
Beyond Static Uncertainty: Modeling Temporal Uncertainty Dynamics for Probabilistic Time Series Forecasting
por: Wang, Yijun, et al.
Publicado: (2026)
por: Wang, Yijun, et al.
Publicado: (2026)
Accelerating MoE Model Inference with Expert Sharding
por: Balmau, Oana, et al.
Publicado: (2025)
por: Balmau, Oana, et al.
Publicado: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
por: Rahman, Musfiqur, et al.
Publicado: (2025)
por: Rahman, Musfiqur, et al.
Publicado: (2025)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
por: Shi, Jiajun, et al.
Publicado: (2025)
por: Shi, Jiajun, et al.
Publicado: (2025)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
por: Chen, Shuhao, et al.
Publicado: (2026)
por: Chen, Shuhao, et al.
Publicado: (2026)
Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments
por: Delgrange, Florent
Publicado: (2026)
por: Delgrange, Florent
Publicado: (2026)
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
por: Verbeken, Brecht, et al.
Publicado: (2026)
por: Verbeken, Brecht, et al.
Publicado: (2026)
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
por: Chen, Simin, et al.
Publicado: (2025)
por: Chen, Simin, et al.
Publicado: (2025)
Noiseless Privacy-Preserving Decentralized Learning
por: Biswas, Sayan, et al.
Publicado: (2024)
por: Biswas, Sayan, et al.
Publicado: (2024)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
por: Zhang, Wenbo, et al.
Publicado: (2025)
por: Zhang, Wenbo, et al.
Publicado: (2025)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
por: Topsakal, Oguzhan, et al.
Publicado: (2024)
por: Topsakal, Oguzhan, et al.
Publicado: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Evaluating Game Difficulty in Tetris Block Puzzle
por: Wang, Chun-Jui, et al.
Publicado: (2026)
por: Wang, Chun-Jui, et al.
Publicado: (2026)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
por: Mayilvahanan, Prasanna, et al.
Publicado: (2025)
por: Mayilvahanan, Prasanna, et al.
Publicado: (2025)
Benchmarking Synthetic Tabular Data: A Multi-Dimensional Evaluation Framework
por: Sidorenko, Andrey, et al.
Publicado: (2025)
por: Sidorenko, Andrey, et al.
Publicado: (2025)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
por: Xu, Yang, et al.
Publicado: (2026)
por: Xu, Yang, et al.
Publicado: (2026)
A Benchmark Environment for Offline Reinforcement Learning in Racing Games
por: Macaluso, Girolamo, et al.
Publicado: (2024)
por: Macaluso, Girolamo, et al.
Publicado: (2024)
Beyond Shapley Values: Cooperative Games for the Interpretation of Machine Learning Models
por: Idrissi, Marouane Il, et al.
Publicado: (2025)
por: Idrissi, Marouane Il, et al.
Publicado: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
por: Novikov, Ivan
Publicado: (2025)
por: Novikov, Ivan
Publicado: (2025)
Monitoring of Static Fairness
por: Henzinger, Thomas A., et al.
Publicado: (2025)
por: Henzinger, Thomas A., et al.
Publicado: (2025)
Leveraging Imperfect Sources to Detect Fairwashing in Black-Box Auditing
por: Bourrée, Jade Garcia, et al.
Publicado: (2023)
por: Bourrée, Jade Garcia, et al.
Publicado: (2023)
The Elicitation Game: Evaluating Capability Elicitation Techniques
por: Hofstätter, Felix, et al.
Publicado: (2025)
por: Hofstätter, Felix, et al.
Publicado: (2025)
Beyond the Meta: Leveraging Game Design Parameters for Patch-Agnostic Esport Analytics
por: Chitayat, Alan Pedrassoli, et al.
Publicado: (2023)
por: Chitayat, Alan Pedrassoli, et al.
Publicado: (2023)
HEARTS: Benchmarking LLM Reasoning on Health Time Series
por: Li, Sirui, et al.
Publicado: (2026)
por: Li, Sirui, et al.
Publicado: (2026)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
por: Li, Yuangang, et al.
Publicado: (2026)
por: Li, Yuangang, et al.
Publicado: (2026)
AuctionNet: A Novel Benchmark for Decision-Making in Large-Scale Games
por: Su, Kefan, et al.
Publicado: (2024)
por: Su, Kefan, et al.
Publicado: (2024)
The Procedural Content Generation Benchmark: An Open-source Testbed for Generative Challenges in Games
por: Khalifa, Ahmed, et al.
Publicado: (2025)
por: Khalifa, Ahmed, et al.
Publicado: (2025)
Ejemplares similares
-
Fairness Auditing with Multi-Agent Collaboration
por: de Vos, Martijn, et al.
Publicado: (2024) -
Optimizing Agentic Workflows using Meta-tools
por: Abuzakuk, Sami, et al.
Publicado: (2026) -
Robust ML Auditing using Prior Knowledge
por: Bourrée, Jade Garcia, et al.
Publicado: (2025) -
Effective LoRA Adapter Routing using Task Representations
por: Dhasade, Akash, et al.
Publicado: (2026) -
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
por: Dorn, Diego, et al.
Publicado: (2024)