Smart Traffic Control
A reinforcement-learning traffic-signal controller trained in a custom four-way intersection simulation to adapt signal phases to changing queue conditions.
Can a learning-based controller outperform a fixed cyclic signal policy?
The project builds a custom Python simulation of a four-way intersection and trains a Q-learning agent to select signal actions from the observed traffic state. The comparison baseline is a Round-Robin controller that cycles through the available phases without learning from congestion.
The optimization target is operational rather than purely algorithmic: increase throughput, reduce average waiting time, and improve the cumulative reward generated by the control policy.
Environment design
- Four-way intersection
- 12 incoming lanes
- Six non-conflicting movement phases
- Poisson vehicle arrivals
- Queue-length observations
- OpenAI Gym-style interface
Discretize the traffic state, then learn which signal action to take.
The state contains queue lengths for all twelve lanes plus the currently active phase. Queue lengths are discretized into congestion categories so a tabular Q-learning policy remains computationally manageable. The agent uses an ε-greedy exploration policy and updates its Q-table from experience.
The first agent was not good enough, and that became part of the engineering result.
Initial model
The early Q-learning configuration performed worse than the Round-Robin baseline and showed high variability. The report attributes this to a coarse reward function, limited state representation, and an inefficient action space.
Reward redesign
The reward function was expanded to penalize waiting and queue buildup more strongly, including severe-congestion penalties, while also rewarding throughput and balanced flow.
State refinement
Queue-length categories were adjusted to distinguish moderate from severe congestion more effectively, improving the agent’s sensitivity to traffic conditions.
Action refinement
The action space was simplified from many unit-by-unit vehicle counts to a smaller predefined set of vehicle batches, improving tractability and interpretability.
After refinement, the learned policy outperformed the fixed baseline across all three tracked metrics.
Total reward
The trained Q-learning agent achieved consistently higher reward than the Round-Robin baseline.
Vehicles passed
Intersection throughput was consistently higher under the learned policy than under the cyclic baseline.
Average waiting time
The trained agent produced substantially shorter waiting times, indicating more responsive congestion management.
Stability
Five independent training runs with different seeds were used to assess variability. The report notes that performance stabilized over time as training progressed.
Environment design mattered as much as the learning algorithm.
The strongest lesson from the project is not simply that “Q-learning works.” The initial learning setup failed to beat a simple baseline. Performance improved only after the state representation, reward function, exploration schedule, and action granularity were redesigned together.
That makes the project a useful example of applied reinforcement learning where system formulation and evaluation design are central to whether the learning algorithm succeeds.
Evaluation metrics
- Total reward
- Vehicles passed
- Average waiting time
- Mean performance over repeated seeds
- Standard deviation across runs
- Visual inspection of simulated traffic flow
A compact RL system with a clear baseline and iterative model redesign.
The project combines simulation, tabular reinforcement learning, operational traffic metrics, and repeated-run validation in one end-to-end experimental system.