Saved in:
| Main Authors: | Marot, Antoine, Rousseau, David, Zhen, Xu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2312.06036 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Managing power grids through topology actions: A comparative study between advanced rule-based and reinforcement learning agents
by: Lehna, Malte, et al.
Published: (2023)
by: Lehna, Malte, et al.
Published: (2023)
Towards certifiable AI in aviation: landscape, challenges, and opportunities
by: Bello, Hymalai, et al.
Published: (2024)
by: Bello, Hymalai, et al.
Published: (2024)
PIQL: Projective Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2025)
by: Han, Xinchen, et al.
Published: (2025)
Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs
by: Sikand, Samarth, et al.
Published: (2025)
by: Sikand, Samarth, et al.
Published: (2025)
Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2026)
by: Han, Xinchen, et al.
Published: (2026)
From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence
by: Theiler, Raffael, et al.
Published: (2026)
by: Theiler, Raffael, et al.
Published: (2026)
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
by: Han, Xinchen, et al.
Published: (2026)
by: Han, Xinchen, et al.
Published: (2026)
How predictable is language model benchmark performance?
by: Owen, David
Published: (2024)
by: Owen, David
Published: (2024)
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations
by: Marchesini, Enrico, et al.
Published: (2025)
by: Marchesini, Enrico, et al.
Published: (2025)
Graph Neural Networks for temporal graphs: State of the art, open challenges, and opportunities
by: Longa, Antonio, et al.
Published: (2023)
by: Longa, Antonio, et al.
Published: (2023)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
Advancing Chinese biomedical text mining with community challenges
by: Zong, Hui, et al.
Published: (2024)
by: Zong, Hui, et al.
Published: (2024)
The challenge of hidden gifts in multi-agent reinforcement learning
by: Malenfant, Dane, et al.
Published: (2025)
by: Malenfant, Dane, et al.
Published: (2025)
Machine Learned Force Fields: Fundamentals, its reach, and challenges
by: Vital, Carlos A., et al.
Published: (2025)
by: Vital, Carlos A., et al.
Published: (2025)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods
by: Clark, Benedict, et al.
Published: (2024)
by: Clark, Benedict, et al.
Published: (2024)
Discovering Data Manifold Geometry via Non-Contracting Flows
by: Vigouroux, David, et al.
Published: (2026)
by: Vigouroux, David, et al.
Published: (2026)
Testing autonomous vehicles and AI: perspectives and challenges from cybersecurity, transparency, robustness and fairness
by: Llorca, David Fernández, et al.
Published: (2024)
by: Llorca, David Fernández, et al.
Published: (2024)
Overcoming classic challenges for artificial neural networks by providing incentives and practice
by: Irie, Kazuki, et al.
Published: (2024)
by: Irie, Kazuki, et al.
Published: (2024)
The role of positional encodings in the ARC benchmark
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
AI Olympics challenge with Evolutionary Soft Actor Critic
by: Calì, Marco, et al.
Published: (2024)
by: Calì, Marco, et al.
Published: (2024)
Opportunities and challenges in the application of large artificial intelligence models in radiology
by: Pan, Liangrui, et al.
Published: (2024)
by: Pan, Liangrui, et al.
Published: (2024)
The OPS-SAT benchmark for detecting anomalies in satellite telemetry
by: Ruszczak, Bogdan, et al.
Published: (2024)
by: Ruszczak, Bogdan, et al.
Published: (2024)
Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications
by: Schouten, Daan, et al.
Published: (2024)
by: Schouten, Daan, et al.
Published: (2024)
The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review
by: Tammisto, Maj-Annika, et al.
Published: (2026)
by: Tammisto, Maj-Annika, et al.
Published: (2026)
Guaranteed prediction sets for functional surrogate models
by: Gray, Ander, et al.
Published: (2025)
by: Gray, Ander, et al.
Published: (2025)
GraphBench: Next-generation graph learning benchmarking
by: Stoll, Timo, et al.
Published: (2025)
by: Stoll, Timo, et al.
Published: (2025)
Robust NAS under adversarial training: benchmark, theory, and beyond
by: Wu, Yongtao, et al.
Published: (2024)
by: Wu, Yongtao, et al.
Published: (2024)
A method to benchmark high-dimensional process drift detection
by: Wolf, Edgar, et al.
Published: (2024)
by: Wolf, Edgar, et al.
Published: (2024)
Cueless EEG imagined speech for subject identification: dataset and benchmarks
by: Derakhshesh, Ali, et al.
Published: (2025)
by: Derakhshesh, Ali, et al.
Published: (2025)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
Interpretability in Symbolic Regression: a benchmark of Explanatory Methods using the Feynman data set
by: Aldeia, Guilherme Seidyo Imai, et al.
Published: (2024)
by: Aldeia, Guilherme Seidyo Imai, et al.
Published: (2024)
CausalRivers -- Scaling up benchmarking of causal discovery for real-world time-series
by: Stein, Gideon, et al.
Published: (2025)
by: Stein, Gideon, et al.
Published: (2025)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
A method for the systematic generation of graph XAI benchmarks via Weisfeiler-Leman coloring
by: Fontanesi, Michele, et al.
Published: (2025)
by: Fontanesi, Michele, et al.
Published: (2025)
Diachronic and synchronic variation in the performance of adaptive machine learning systems: The ethical challenges
by: Hatherley, Joshua, et al.
Published: (2025)
by: Hatherley, Joshua, et al.
Published: (2025)
The impact of internal variability on benchmarking deep learning climate emulators
by: Lütjens, Björn, et al.
Published: (2024)
by: Lütjens, Björn, et al.
Published: (2024)
Application of predictive machine learning in pen & paper RPG game design
by: Śliwa, Jolanta
Published: (2025)
by: Śliwa, Jolanta
Published: (2025)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Similar Items
-
Managing power grids through topology actions: A comparative study between advanced rule-based and reinforcement learning agents
by: Lehna, Malte, et al.
Published: (2023) -
Towards certifiable AI in aviation: landscape, challenges, and opportunities
by: Bello, Hymalai, et al.
Published: (2024) -
PIQL: Projective Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2025) -
Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs
by: Sikand, Samarth, et al.
Published: (2025) -
Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2026)