Welcome to nature.com. We appreciate your visit. If you are experiencing limitations in viewing due to your current browser version with limited CSS support, we recommend using a more updated browser for the best experience. Alternatively, switching off compatibility mode in Internet Explorer may also help. While we work to ensure continued support, please note that the site is displayed without styles and JavaScript.
This article from Scientific Reports (2026) is being shared in advance to offer expedited access to peer-reviewed research. It is citable and includes a permanent DOI. Please be aware that this version is subject to further revisions and will eventually be replaced by the final Version of Record. All standard legal disclaimers apply.
This study focuses on the exploration strategy in deep reinforcement learning for robotic manipulation. The research compares three off-policy configurations for a 6-degree-of-freedom robotic grasping task using a Universal Robots UR5e in the RoboSuite simulation environment. The configurations studied are Soft Actor-Critic (SAC) with automatic entropy tuning, Twin Delayed Deep Deterministic Policy Gradient (TD3) with piecewise linear noise decay, and TD3 with constant Gaussian noise. They all have identical network architectures, reward functions, and training budgets.
Throughout 10,000 episodes across five independent runs, each configuration was trained and evaluated over 1,000 deterministic episodes within and beyond the training distribution. SAC demonstrated the highest success rate (89.6%) within the training distribution, along with superior training efficiency compared to both TD3 variants, which were statistically similar within the training distribution. However, under out-of-distribution evaluation, the noise-decay variant outperformed the constant noise variant significantly.
The study emphasizes that the design of noise schedules has a notable impact on both in-distribution performance and spatial out-of-distribution generalization. The research suggests that standard evaluation protocols may not effectively capture the differences between exploration strategies. The authors acknowledge Prof. Dushko Stavrov for his contributions and are affiliated with the Faculty of Electrical Engineering and Information Technologies, Ss. Cyril and Methodius University in Skopje, North Macedonia.
This article is licensed under a Creative Commons Attribution 4.0 International License, allowing for sharing, adaptation, and distribution with appropriate credit. For further details on usage permissions, please refer to the full license available at http://creativecommons.org/licenses/by/4.0/. For citation and additional information, please access the article through the provided DOI link.
For more insights on AI and Robotics research, consider signing up for the Nature Briefing: AI and Robotics newsletter, delivering the latest updates weekly to your inbox.
