Deep reinforcement learning has achieved substantial performance gains over classical control approaches.
Yet, a central challenge to learning in real-world applications is acquiring costly samples.
Kolmogorov-Arnold Networks are a recently proposed architecture that can learn physical relationships in control problems effectively, with significantly higher parameter efficiency and interpretability when compared to Multi-Layer-Perceptron architectures.
In this work
We systematically study sample-efficiency using computational experiments, covering the Feynman dataset and the Gymnasium RL benchmark.
The results show that similar performance can be achieved with 40% fewer samples using the Kolmogorov-Arnold architecture, and that relative performance improvements up to 50% occur during the training process.
The observed gains are robust to varying levels of noise in rewards.
These results highlight the potential of the Kolmogorov-Arnold architectures for more sample-efficient reinforcement learning.
Code: https://github.com/DerKevinRiehl/neurips26_kan_training