Saved in:
Bibliographic Details
Main Author: Bereketoglu, Abdullah Burkan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.06323
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909641456222208
author Bereketoglu, Abdullah Burkan
author_facet Bereketoglu, Abdullah Burkan
contents Model-free and reinforcement learning-based adaptive filtering methods are gaining traction for denoising in dynamic, non-stationary environments such as wireless signal channels. Traditional filters like LMS, RLS, Wiener, and Kalman are limited by assumptions of stationary or requiring complex fine-tuning or exact noise statistics or fixed models. This letter proposes an adaptive filtering framework using Proximal Policy Optimization (PPO), guided by a composite reward that balances SNR improvement, MSE reduction, and residual smoothness. Experiments on synthetic signals with various noise types show that our PPO agent generalizes beyond its training distribution, achieving real-time performance and outperforming classical filters. This work demonstrates the viability of policy-gradient reinforcement learning for robust, low-latency adaptive signal filtering.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Composite Reward Design in PPO-Driven Adaptive Filtering
Bereketoglu, Abdullah Burkan
Signal Processing
Machine Learning
Systems and Control
Model-free and reinforcement learning-based adaptive filtering methods are gaining traction for denoising in dynamic, non-stationary environments such as wireless signal channels. Traditional filters like LMS, RLS, Wiener, and Kalman are limited by assumptions of stationary or requiring complex fine-tuning or exact noise statistics or fixed models. This letter proposes an adaptive filtering framework using Proximal Policy Optimization (PPO), guided by a composite reward that balances SNR improvement, MSE reduction, and residual smoothness. Experiments on synthetic signals with various noise types show that our PPO agent generalizes beyond its training distribution, achieving real-time performance and outperforming classical filters. This work demonstrates the viability of policy-gradient reinforcement learning for robust, low-latency adaptive signal filtering.
title Composite Reward Design in PPO-Driven Adaptive Filtering
topic Signal Processing
Machine Learning
Systems and Control
url https://arxiv.org/abs/2506.06323