Date Approved
7-20-2026
Embargo Period
7-20-2026
Document Type
Thesis
Degree Name
M.S. Data Science
Department
Computer Science
College
College of Science & Mathematics
Advisor
Huaxia Wang, Ph.D.
Committee Member 1
Ben Wu, Ph.D.
Committee Member 2
Gianesh Baliga, Ph.D.
Committee Member 3
Silvija Kokalj-Filipovic, Ph.D.
Committee Member 4
Xiajun Jiang, Ph.D.
Keywords
Deep Reinforcement Learning;Integrated Sensing and Communication;Optical Wireless Communication;Vehicle-to-Everything
Abstract
Optical Wireless Communication (OWC) promises ultra-high bandwidth for vehicular networks but suffers from critical beam misalignment in dynamic environments. The emerging paradigm of Optical Integrated Sensing and Communication (O-ISAC) addresses this challenge by unifying the sensing and communication functions over a single shared optical aperture, enabling real-time environmental awareness to guide beam steering. However, the multi-objective, stochastic nature of this joint optimization problem renders classical closed-form controllers inadequate. This thesis proposes and evaluates a deep reinforcement learning framework for continuous beam tracking in a 2D vehicular O-ISAC scenario. We conduct a comprehensive comparative study of four state-of-the-art deep RL algorithms: Soft Actor-Critic (SAC), Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO), and Deep Deterministic Policy Gradient (DDPG), each trained for 5 million environment steps in a custom Gymnasium simulation. The environment incorporates log-normal atmospheric turbulence, geometric pointing loss, ray-casting blockage detection, and stochastic vehicle dynamics for both vehicle-to-everything (V2V) and vehicle-to-infrastructure (V2I) links. Our results demonstrate that SAC, through entropy-regularized maximum-entropy reinforcement learning, achieves superior and robust beam tracking performance. SAC attains a mean evaluation reward of 47,350 versus 31,200 for DDPG, 18,900 for PPO, and 12,800 for TRPO, and maintains simultaneous dual-link connectivity 78.3% of the time compared to 54.1% for the myopic greedy baseline. Temporal analysis confirms that SAC's learned policy sustains link quality 1.45 times more effectively than the reactive baseline during dynamic intersection crossing scenarios. These results establish O-ISAC as a compelling architecture for next-generation vehicular networks and deep RL as the methodology of choice for joint beam management.
Recommended Citation
Sharan, Akshat, "OPTICAL INTEGRATED SENSING AND COMMUNICATION: DEEP REINFORCEMENT LEARNING-BASED BEAM TRACKING FOR DYNAMIC VEHICLE-TO-EVERYTHING NETWORKS" (2026). Theses and Dissertations. 3564.
https://rdw.rowan.edu/etd/3564