Bio-inspired aerial robots, which mimic flight techniques observed in birds and insects, offer distinct advantages over traditional UAVs in challenging maneuvers like perching. During perching, the vehicle is required to approach a target and reduce most of its kinetic energy within a short distance. The control of these systems is intricate due to their nonlinear dynamics, strong connection between translation and rotation, and aerodynamic variations caused by their variable wing and tail structures.
Conventional control methods based on linear models have limitations in such scenarios, prompting the exploration of alternative strategies based on reinforcement learning. The primary objective of this Master’s Thesis is to create and assess a modeling and control system for a bio-inspired aerial robot utilizing deep reinforcement learning to analyze its capability in executing approach and perching maneuvers. The study is built upon the foundational research by Wüest et al., employing a simplified two-dimensional longitudinal model to focus on key variables: position, velocity, pitch attitude, and angle of attack.
A simulation environment compatible with Gymnasium and Stable-Baselines3 was established to define the observation space, action space, reward function, and episode termination criteria. Two reinforcement learning algorithms, Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC), were trained and compared within this framework. Additionally, different success margins’ impacts on the learned control behavior were examined.
The outcomes reveal that both algorithms successfully learn control policies for perching, demonstrating coherent behaviors like decreased forward velocity, increased pitch angle and angle of attack in the final phase, and coordinated actions between thrust, elevator deflection, and variable wing and tail morphology. PPO exhibits consistent performance for various success margins, while SAC displays enhanced robustness, generating smoother trajectories under stricter conditions. However, a limitation is recognized where the final vehicle velocity remains around 4 m/s in most scenarios, depicting a controlled approach rather than a complete stop during perching.
In conclusion, this study showcases that deep reinforcement learning serves as a flexible and viable framework for investigating approach and perching strategies in bio-inspired aerial robots. Furthermore, it sets the groundwork for future advancements towards three-dimensional models, improved reward systems, and potential implementation on real robotic platforms.
