Probability Hacking and the Design of Trustworthy ML for Signal Processing in C-UAS: A Scenario Based Method
Fuente:
arXiv
Saved in:
| Main Authors: | Janssens, Liisa, Middeldorp, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Benchmark Hacking in ML Contests: Modeling, Insights and Design
by: Qiu, Xiaoyun, et al.
Published: (2026)
by: Qiu, Xiaoyun, et al.
Published: (2026)
X Hacking: The Threat of Misguided AutoML
by: Sharma, Rahul, et al.
Published: (2024)
by: Sharma, Rahul, et al.
Published: (2024)
Enhancing Trustworthiness in ML-Based Network Intrusion Detection with Uncertainty Quantification
by: Talpini, Jacopo, et al.
Published: (2023)
by: Talpini, Jacopo, et al.
Published: (2023)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
by: Wu, Rui, et al.
Published: (2026)
by: Wu, Rui, et al.
Published: (2026)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Beyond the Single-Best Model: Rashomon Partial Dependence Profile for Trustworthy Explanations in AutoML
by: Cavus, Mustafa, et al.
Published: (2025)
by: Cavus, Mustafa, et al.
Published: (2025)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
by: Binkyte, Ruta, et al.
Published: (2025)
by: Binkyte, Ruta, et al.
Published: (2025)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
by: Li, Lichen, et al.
Published: (2026)
by: Li, Lichen, et al.
Published: (2026)
Estimating the Joint Probability of Scenario Parameters with Gaussian Mixture Copula Models
by: Reichenbächer, Christian, et al.
Published: (2025)
by: Reichenbächer, Christian, et al.
Published: (2025)
Hacking Task Confounder in Meta-Learning
by: Wang, Jingyao, et al.
Published: (2023)
by: Wang, Jingyao, et al.
Published: (2023)
A Trustworthy By Design Classification Model for Building Energy Retrofit Decision Support
by: Rempi, Panagiota, et al.
Published: (2025)
by: Rempi, Panagiota, et al.
Published: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025)
by: Khalaf, Hadi, et al.
Published: (2025)
FedDRL: A Trustworthy Federated Learning Model Fusion Method Based on Staged Reinforcement Learning
by: Chen, Leiming, et al.
Published: (2023)
by: Chen, Leiming, et al.
Published: (2023)
RL-based Control of UAS Subject to Significant Disturbance
by: Chakraborty, Kousheek, et al.
Published: (2025)
by: Chakraborty, Kousheek, et al.
Published: (2025)
When Small Models Are Right for Wrong Reasons: Process Verification for Trustworthy Agents
by: Advani, Laksh
Published: (2026)
by: Advani, Laksh
Published: (2026)
Trustworthy Graph Neural Networks: Aspects, Methods and Trends
by: Zhang, He, et al.
Published: (2022)
by: Zhang, He, et al.
Published: (2022)
Hydro: Adaptive Query Processing of ML Queries
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
One Model Many Scores: Using Multiverse Analysis to Prevent Fairness Hacking and Evaluate the Influence of Model Design Decisions
by: Simson, Jan, et al.
Published: (2023)
by: Simson, Jan, et al.
Published: (2023)
EvilGenie: A Reward Hacking Benchmark
by: Gabor, Jonathan, et al.
Published: (2025)
by: Gabor, Jonathan, et al.
Published: (2025)
Trustworthy Prediction with Gaussian Process Knowledge Scores
by: Butler, Kurt, et al.
Published: (2025)
by: Butler, Kurt, et al.
Published: (2025)
Accelerating RF Power Amplifier Design via Intelligent Sampling and ML-Based Parameter Tuning
by: Sriram, Abhishek, et al.
Published: (2025)
by: Sriram, Abhishek, et al.
Published: (2025)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
Using Quality Attribute Scenarios for ML Model Test Case Generation
by: Brower-Sinning, Rachel, et al.
Published: (2024)
by: Brower-Sinning, Rachel, et al.
Published: (2024)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
by: Wang, Songtao, et al.
Published: (2026)
by: Wang, Songtao, et al.
Published: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Trustworthy Classification through Rank-Based Conformal Prediction Sets
by: Luo, Rui, et al.
Published: (2024)
by: Luo, Rui, et al.
Published: (2024)
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
A Framework to Model ML Engineering Processes
by: Morales, Sergio, et al.
Published: (2024)
by: Morales, Sergio, et al.
Published: (2024)
Elucidating the Design Choice of Probability Paths in Flow Matching for Forecasting
by: Lim, Soon Hoe, et al.
Published: (2024)
by: Lim, Soon Hoe, et al.
Published: (2024)
A Method For Bounding Tail Probabilities
by: Zlatanov, Nikola
Published: (2024)
by: Zlatanov, Nikola
Published: (2024)
Trustworthy Transfer Learning: A Survey
by: Wu, Jun, et al.
Published: (2024)
by: Wu, Jun, et al.
Published: (2024)
OpenUAS: Embeddings of Cities in Japan with Anchor Data for Cross-city Analysis of Area Usage Patterns
by: Tamura, Naoki, et al.
Published: (2024)
by: Tamura, Naoki, et al.
Published: (2024)
Missing Data in Signal Processing and Machine Learning: Models, Methods and Modern Approaches
by: Hippert-Ferrer, Alexandre, et al.
Published: (2025)
by: Hippert-Ferrer, Alexandre, et al.
Published: (2025)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Similar Items
-
On Benchmark Hacking in ML Contests: Modeling, Insights and Design
by: Qiu, Xiaoyun, et al.
Published: (2026) -
X Hacking: The Threat of Misguided AutoML
by: Sharma, Rahul, et al.
Published: (2024) -
Enhancing Trustworthiness in ML-Based Network Intrusion Detection with Uncertainty Quantification
by: Talpini, Jacopo, et al.
Published: (2023) -
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
by: Wu, Rui, et al.
Published: (2026) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)