UTAR Institutional Repository

AI-driven adaptive traffic light optimization using dueling deep Q-networks

Lau, Yee Hang (2026) AI-driven adaptive traffic light optimization using dueling deep Q-networks. Final Year Project, UTAR.

[img] PDF
Download (2024Kb)

    Abstract

    This study presents a Dueling Double Deep Q-Network (Dueling DDQN) based adaptive traffic signal control framework for a signalised four-way intersection simulated in SUMO (Simulation of Urban Mobility). Unlike traditional fixed-time controllers, the proposed agent dynamically selects among three discrete signal phases based on a 13-dimensional state vector encoding one-hot phase representation, directional vehicle queues and waiting times, and pedestrian demand. The reward function is a seven-component composite that balances vehicle efficiency through queue reduction and throughput bonuses, safety through a flat red-light violation penalty, and pedestrian service through a fixed penalty weight. A controlled ablation study compares four DQN architectures — VanillaDQN, DoubleDQN, DuelingDQN, and DuelingDoubleDQN — under identical reward functions, hyperparameters, and training episodes across three demand scenarios (low: ~905 veh/hr, medium: ~1,809 veh/hr, high: ~3,165 veh/hr). All models are trained on the medium scenario only, and generalisation is evaluated by testing on unseen low and high demand scenarios without retraining. Results demonstrate that DuelingDoubleDQN achieves a 54.1% reduction in average vehicle waiting time and a 49.3% reduction in vehicle queue length on the training scenario compared to VanillaDQN. Under the unseen high-demand scenario, DuelingDoubleDQN reduces vehicle wait by 21.5%, confirming policy generalisation beyond the training distribution. A key ablation finding is that DuelingDQN without the double correction performs worst at high demand (1,095.98s wait), exceeding even VanillaDQN, which confirms that the double update rule is critical for stable performance This study presents a Dueling Double Deep Q-Network (Dueling DDQN) based adaptive traffic signal control framework for a signalised four-way intersection simulated in SUMO (Simulation of Urban Mobility). Unlike traditional fixed-time controllers, the proposed agent dynamically selects among three discrete signal phases based on a 13-dimensional state vector encoding one-hot phase representation, directional vehicle queues and waiting times, and pedestrian demand. The reward function is a seven-component composite that balances vehicle efficiency through queue reduction and throughput bonuses, safety through a flat red-light violation penalty, and pedestrian service through a fixed penalty weight. A controlled ablation study compares four DQN architectures — VanillaDQN, DoubleDQN, DuelingDQN, and DuelingDoubleDQN — under identical reward functions, hyperparameters, and training episodes across three demand scenarios (low: ~905 veh/hr, medium: ~1,809 veh/hr, high: ~3,165 veh/hr). All models are trained on the medium scenario only, and generalisation is evaluated by testing on unseen low and high demand scenarios without retraining. Results demonstrate that DuelingDoubleDQN achieves a 54.1% reduction in average vehicle waiting time and a 49.3% reduction in vehicle queue length on the training scenario compared to VanillaDQN. Under the unseen high-demand scenario, DuelingDoubleDQN reduces vehicle wait by 21.5%, confirming policy generalisation beyond the training distribution. A key ablation finding is that DuelingDQN without the double correction performs worst at high demand (1,095.98s wait), exceeding even VanillaDQN, which confirms that the double update rule is critical for stable performance under near-saturation conditions. An automated red-light violation monitoring module detects, classifies, and logs each violation event with timestamp, vehicle ID, approach direction, speed, and phase state, enabling precision/recall/F1 evaluation. A reward design limitation regarding pedestrian service deprioritisation is identified and investigated through a constrained reward variant. Findings contribute towards safer, more adaptive, and computationally tractable intelligent transportation systems, under near-saturation conditions. An automated red-light violation monitoring module detects, classifies, and logs each violation event with timestamp, vehicle ID, approach direction, speed, and phase state, enabling precision/recall/F1 evaluation. A reward design limitation regarding pedestrian service deprioritisation is identified and investigated through a constrained reward variant. Findings contribute towards safer, more adaptive, and computationally tractable intelligent transportation systems.

    Item Type: Final Year Project / Dissertation / Thesis (Final Year Project)
    Subjects: T Technology > T Technology (General)
    Divisions: Faculty of Information and Communication Technology > Bachelor of Computer Science (Honours)
    Depositing User: ML Main Library
    Date Deposited: 06 Aug 2026 19:24
    Last Modified: 06 Aug 2026 19:24
    URI: http://eprints.utar.edu.my/id/eprint/7763

    Actions (login required)

    View Item