Welcome to nature.com! We appreciate your visit, but it seems you are using a browser version that has limited support for CSS. For the best experience, we recommend updating your browser or disabling compatibility mode in Internet Explorer. While we work on ensuring continued support, the site is being displayed without styles and JavaScript.
Scientific Reports (2026) acknowledges 981 accesses to this article, with 1 Altmetric mentioning its details. We are sharing an unedited version of this manuscript to grant early access to its discoveries. Please be aware that prior to final publication, the manuscript will undergo further editing, and errors may exist that might affect the content, subject to all legal disclaimers.
The utilization of robots has increased significantly across various sectors such as household applications, surveillance, medical services, manufacturing, and logistics automation. Developing dependable control mechanisms that can adapt to changing environments is imperative to ensure the efficient and safe operation of these robots.
Reinforcement learning (RL), a technique situated between supervised and unsupervised learning, deals with learning in sequential decision-making scenarios where feedback is limited. This paper introduces a structured framework and comprehensive overview of the Markov Decision Process (MDP), including explanations of value functions and policies. The article elucidates the principles and concepts of MDPs and discusses RL algorithms that compute optimal behaviors using dynamic programming based on Q-learning.
Robotics path planning environments are often modeled as MDPs under RL, aiming to learn a control strategy that maximizes the total reward. This paper primarily focuses on introducing fundamental MDP-based algorithm information to achieve optimal behaviors, exploring various interpretations of optimality in decision-making sequences.
Furthermore, a simulation experiment validates the theoretical aspects outlined in the study, employing the Q-learning approach in a custom robot environment with diverse obstacles positioned at different locations. The simulation demonstrates an accuracy rate of 88.80% for the robot to reach the designated target point and includes a comparative performance analysis between Q-learning and DQN algorithms.
This research received financial support from the Excellence Initiative – Research University – AGH, Action D11, Project No. 15598 of the AGH University of Kraków, Poland. The authors, Ravi Raj of the Faculty of Computer Science, Electronics, and Telecommunications at AGH University of Krakow, and Andrzej Kos, declare no conflicting interests.
This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, permitting non-commercial use, sharing, distribution, and reproduction with proper credit to the original author(s) and source. Access to adapted material derived from this article requires direct permission from the copyright holder.
For more information on the Markov decision process for reinforcement learning in robotics perception, refer to the publication in Scientific Reports (2026). Stay updated with the latest in AI and robotics research by signing up for the Nature Briefing: AI and Robotics newsletter, delivered to your inbox weekly.
Thank you for your interest and contribution to the scientific community.
