Abstract:Most safe reinforcement learning (RL) algorithms depend on the accurate reward that is rarely available in wireless security applications and suffer from severe performance degradation for the learning agents that have to choose the policy from a large action set. In this paper, we propose a safe RL algorithm, which uses a policy priority-based hierarchical structure to divide each policy into sub-policies with different selection priorities and thus compresses the action set. By applying inter-agent transfer learning to initialize the learning parameters, this algorithm accelerates the initial exploration of the optimal policy. Based on a security criterion that evaluates the risk value, the sub-policy distribution formulation avoids the dangerous sub-policies that cause learning failure such as severe network security problems in wireless security applications, e.g., Internet services interruption. We also propose a deep safe RL and design four deep neural networks in each sub-policy selection to further improve the learning efficiency for the learning agents that support four convolutional neural networks (CNNs): The Q-network evaluates the long-term expected reward of each sub-policy under the current state, and the E-network evaluates the long-term risk value. The target Q and E-networks update the learning parameters of the corresponding CNN to improve the policy exploration stability. As a case study, our proposed safe RL algorithms are implemented in the anti-jamming communication of unmanned aerial vehicles (UAVs) to select the frequency channel and transmit power to the ground node. Experimental results show that our proposed schemes significantly improve the UAV communication performance, save the UAV energy and increase the reward compared with the benchmark against jamming.

UAV Networks Against Multiple Maneuvering Smart Jamming with Knowledge-Based Reinforcement Learning.

Multi-Agent Reinforcement Learning Based UAV Swarm Communications Against Jamming

UAV Communication Against Intelligent Jamming: A Stackelberg Game Approach With Federated Reinforcement Learning

Penalized Reinforcement Learning-Based Energy-Efficient UAV-RIS Assisted Maritime Uplink Communications Against Jamming

UAV Anti-Jamming Video Transmissions with QoE Guarantee: A Reinforcement Learning-Based Approach.

Reinforcement Learning-Based Anti-Jamming in Networked UAV Radar Systems

Jamming-Resilient Path Planning for Multiple UAVs via Deep Reinforcement Learning

A Learning Approach Towards Secure Cognitive Networks with UAV Relaying and Active Jamming

Safe Exploration in Wireless Security: A Safe Reinforcement Learning Algorithm With Hierarchical Structure

Smart Shield: Prevent Aerial Eavesdropping Via Cooperative Intelligent Jamming Based on Multi-Agent Reinforcement Learning

Energy Efficient Relay in UAV Networks Against Jamming: A Reinforcement Learning Based Approach.

Distributed reinforcement learning based framework for energy-efficient UAV relay against jamming

Lightweight Reinforcement Learning with State Abstraction for Dynamic Spectrum Anti-Jamming Communications

Knowledge Aided Model-Based Reinforcement Learning for Anti-Jamming Strategy Learning

Collaborative Multiagent Reinforcement Learning Aided Resource Allocation for UAV Anti-Jamming Communication

UAV-Enabled Secure Communications by Multi-Agent Deep Reinforcement Learning

Cooperative Multi-UAV Dynamic Anti-Jamming Scheme with Deep Reinforcement Learning.

Reinforcement Learning Based Energy Efficient Robot Relay for Unmanned Aerial Vehicles Against Smart Jamming

Deep Reinforcement Learning-Driven Jamming-Enhanced Secure Unmanned Aerial Vehicle Communications

Meta-Reinforcement Learning in Time-Varying UAV Communications: Adaptive Anti-Jamming Channel Selection

Resilient Path Planning for UAVs in Data Collection under Adversarial Attacks