Projects / Paper

Centralized PPO-Based DRL for Multi-UAV-BS Positioning and Trajectory Optimization in Disaster Response Networks

Akhtarshenas, A., Ibáñez, M. R., Bernabé, M., López-Pérez, D., Debbah, M.
Submitted to IEEE Transactions on Machine Learning in Communications and Networking

Unmanned aerial vehicle-mounted base stations (UAV-BSs) offer a flexible way to restore connectivity in GPS-free emergency scenarios, where rapid deployment of communication infrastructure is critical for search-and-rescue operations. This work extends a centralized learning framework to a multi-UAV-BS architecture, where a central UAV acts as an intelligent agent coordinating the 3D positioning of multiple serving UAV-BSs.

The problem is formulated as a fairness-aware sum-throughput maximization task and solved with Proximal Policy Optimization (PPO). The agent relies only on radio-sensing measurements — received power, SINR, and angle-of-arrival — with no GPS information from ground users. Across static, linear, circular, cosine, and composite mobility patterns, the proposed PPO framework consistently outperforms DQN and DDPG in convergence stability, mean reward, and throughput.