Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<jats:p>Learning from Demonstration (LfD) is an attractive way to transfer manipulation skills to industrial robots, but it is inherently imitative: the learnt skill is only as good as the trajectory it reproduces. Reinforcement learning (RL) can refine such a skill, yet suffers from high sample cost and the difficulty of hand-designing a single scalar reward. This paper presents a three-stage refinement framework coupling Dynamic Movement Primitives (DMPs) with deep RL, evaluated with prescribed analytic reference trajectories standing in for demonstrations; the pipeline is agnostic to the demonstration source. The DMP forcing term is recast as a policy network with a position/velocity-enriched state and a spatial-scaling reformulation that preserves classical-DMP generalisation. The policy is bootstrapped by supervised learning and refined with actor–critic deep RL (TD3, SAC, PPO). A prediction-guided multi-objective stage builds a Pareto front trading imitation fidelity against motion effort, and a TOPSIS step selects a balanced operating point without manual reward weighting, once the objectives share a common scale. The framework is validated on a pick-and-place task in MuJoCo and on a physical manipulator in three experiments totalling 850 trials. The selected policy lowers the integrated transport speed by 39.2% relative to the imitation-fidelity extreme of its Pareto front, at the highest task-success rate (250 trials); a condition-coverage variant transfers to unseen start, goal, payload, friction, height and obstacle variations with a 0.6-point success gap (360 trials); and the balanced policy cuts dynamic actuator effort, jerk and vibration by tens of percent at non-inferior accuracy (240 trials).</jats:p>