Supervised Project · SCE · Computer Science · Beer Sheva

Smart Traffic Control

A reinforcement-learning traffic-signal controller trained in a custom four-way intersection simulation to adapt signal phases to changing queue conditions.

2024–2025Or Avital Butbul · Lotem CohenQ-learningReinforcement learningTraffic simulationAdaptive control
12incoming lanes
6signal phases
6,000training episodes per reported run
5independent seeds in stability analysis
Research question

Can a learning-based controller outperform a fixed cyclic signal policy?

The project builds a custom Python simulation of a four-way intersection and trains a Q-learning agent to select signal actions from the observed traffic state. The comparison baseline is a Round-Robin controller that cycles through the available phases without learning from congestion.

The optimization target is operational rather than purely algorithmic: increase throughput, reduce average waiting time, and improve the cumulative reward generated by the control policy.

Environment design

  • Four-way intersection
  • 12 incoming lanes
  • Six non-conflicting movement phases
  • Poisson vehicle arrivals
  • Queue-length observations
  • OpenAI Gym-style interface
Learning pipeline

Discretize the traffic state, then learn which signal action to take.

The state contains queue lengths for all twelve lanes plus the currently active phase. Queue lengths are discretized into congestion categories so a tabular Q-learning policy remains computationally manageable. The agent uses an ε-greedy exploration policy and updates its Q-table from experience.

Traffic-control pipeline from four-way intersection simulation through discretized state representation and Q-learning to adaptive signal actions, evaluated against a Round-Robin baseline.
The controller is evaluated on reward, throughput, and average waiting time rather than on reward alone.
What changed during development

The first agent was not good enough, and that became part of the engineering result.

Initial model

The early Q-learning configuration performed worse than the Round-Robin baseline and showed high variability. The report attributes this to a coarse reward function, limited state representation, and an inefficient action space.

Reward redesign

The reward function was expanded to penalize waiting and queue buildup more strongly, including severe-congestion penalties, while also rewarding throughput and balanced flow.

State refinement

Queue-length categories were adjusted to distinguish moderate from severe congestion more effectively, improving the agent’s sensitivity to traffic conditions.

Action refinement

The action space was simplified from many unit-by-unit vehicle counts to a smaller predefined set of vehicle batches, improving tractability and interpretability.

Reported results

After refinement, the learned policy outperformed the fixed baseline across all three tracked metrics.

Total reward

The trained Q-learning agent achieved consistently higher reward than the Round-Robin baseline.

Vehicles passed

Intersection throughput was consistently higher under the learned policy than under the cyclic baseline.

Average waiting time

The trained agent produced substantially shorter waiting times, indicating more responsive congestion management.

Stability

Five independent training runs with different seeds were used to assess variability. The report notes that performance stabilized over time as training progressed.

Interpretation

Environment design mattered as much as the learning algorithm.

The strongest lesson from the project is not simply that “Q-learning works.” The initial learning setup failed to beat a simple baseline. Performance improved only after the state representation, reward function, exploration schedule, and action granularity were redesigned together.

That makes the project a useful example of applied reinforcement learning where system formulation and evaluation design are central to whether the learning algorithm succeeds.

Evaluation metrics

  • Total reward
  • Vehicles passed
  • Average waiting time
  • Mean performance over repeated seeds
  • Standard deviation across runs
  • Visual inspection of simulated traffic flow
Project materials

Final report.

The student final report for this project, as a PDF.

Supervised project

A compact RL system with a clear baseline and iterative model redesign.

The project combines simulation, tabular reinforcement learning, operational traffic metrics, and repeated-run validation in one end-to-end experimental system.