DOI: 10.3390/info17090911 ISSN: 2078-2489

A Review of Joint Unmanned Aerial Vehicle Trajectory and Camera Orientation Optimization

Jakub Kůdela

Camera-equipped Unmanned Aerial Vehicle (UAV) planning couples vehicle motion, camera pose, and scene-dependent sensing utility. The relevant literature is distributed across aerial reconstruction, inspection, target tracking, active perception, cinematography, and coverage planning, with substantial differences in vehicle models, camera mechanisms, visibility assumptions, and evaluation practice. A common vehicle–camera formulation is used here to compare two physical camera-realization mechanisms—independent gimbal actuation and body-coupled orientation—while treating viewpoint-first camera-pose planning as a separate representation that may defer physical realization. Mixed-integer formulations are first examined in detail because coverage, visibility, sequencing, assignment, and discrete camera modes introduce a logical structure; a representative time-expanded MILP is then provided for joint motion–view selection. Evolutionary methods are reviewed from early genetic and differential-evolution UAV planners through information-driven, constrained multiobjective, and hybrid formulations; and cooperative, surrogate-assisted, and transferability-aware methods from adjacent problem classes are then examined as possible extensions. In the frozen coded corpus, 13 direct studies use a physically independently actuated camera/gimbal, but none uses an evolutionary method as the primary optimizer for joint vehicle–gimbal motion; direct evolutionary joint vehicle–gimbal optimization therefore remains sparse. Simulation-to-reality transfer is analyzed as mismatch in dynamics, tracking, gimbal response, calibration, image formation, scene geometry, perception, and timing, with corresponding discussion of randomization, adaptive models, HIL evaluation, robust optimization, and transferability-aware search. A structured reproducibility audit through 31 August 2026 freezes the evidence base at 124 coded sources (62 direct studies, 48 adjacent precedents, and 14 proposed-transfer sources) and confirms persistent fragmentation in scenes, sensors, metrics, and computational budgets. The remaining technical questions concern mechanism-aware camera realization, scalable visibility, multi-view sensing utility under feedback, solver decomposition for mixed discrete–continuous problems, and transfer-sensitive evaluation. By synthesizing these methodological differences, this review identifies persistent research gaps and formulates recommendations for algorithm design, hybrid optimization, benchmarking, and sim-to-real validation.