Abstract
<title>Abstract</title> <p>This paper addresses the trajectory-tracking control problem of multi-DOF robotic manipulators operating under significant payload variations. Existing adaptive and optimized PID/time-delay control methods mainly tune or adapt controller gains while retaining their underlying control laws and have primarily been validated on low-DOF manipulators under simple payload profiles, limiting their ability to maintain accurate and smooth tracking under complex time-varying payload conditions. To address these limitations, this paper develops a model-free deep reinforcement learning (DRL) controller based on soft actor-critic algorithm with payload domain randomization. The trajectory-tracking problem is formulated as a Markov decision process with an augmented state vector comprising three groups: current joint positions and velocities together with their corresponding tracking errors; past joint positions, velocities, and torque commands; and future reference joint positions and velocities. A tracking-oriented reward function is then designed to balance tracking accuracy, torque efficiency, and control smoothness. The proposed controller is evaluated through simulation results on a commercial 6-DOF UR5e manipulator in MuJoCo under a high-speed sinusoidal trajectory and four representative payload scenarios: constant, linearly decreasing, abrupt-changing, and complex time-varying payloads. In the most challenging complex time-varying payload scenario, where the payload varies from 5 kg to 0 kg through constant phases, gradual variations, and abrupt changes, the proposed controller decreases the average absolute error from 19.41o to 1.41o , root mean square error from 27.38o to 1.75o, average absolute torque from 49.22 Nm to 22.25 Nm, and average absolute variation of torque from 5.26Nm to 1.72 Nm relative to the PID controller. These results represent approximate improvements of 92.72%, 93.6%, 54.78%, and 67.24%, respectively, demonstrating improved tracking accuracy, transient response, control effort, and torque smoothness under challenging payload-varying conditions.</p>