Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Urban rail operators must adapt full-length and short-turn services to time-varying passenger demand while satisfying operational constraints. We formulate this task as a constrained Markov decision process and develop an action-masked deep Q-network (DQN) that excludes infeasible decisions according to headway limits, turn-back capacity, rolling-stock availability and the minimum interval between service-pattern changes. In simulations of a 12-station line under a typical weekday demand profile, the proposed method achieved a mean reward of 93.22. Relative to fixed full-length, fixed short-turn and threshold-based policies, it reduced the waiting-time metric by 83.12%, 62.78% and 76.84%, respectively, while increasing the number of passengers served by 19.42%, 17.39% and 5.08%. The action mask reduced the invalid-action rate during training from 0.2274 to 0, although the ablation results did not show a corresponding improvement in reward or waiting time over an unmasked DQN. These results indicate that action masking can enforce operational feasibility while enabling demand-responsive service selection, but its effects on policy performance depend on the training configuration and operating scenario.</p>

Show More

Keywords

while fulllength shortturn demand operational

Related Articles

PORE

About

Connect