Abstract
<jats:p>Action-selection is determined by a combination of goal-directed and habitual processes. Habits are defined as the reward-independent, stimulus-response relationships which form when an action is regularly executed in the same context, regardless of outcome. An influential computational model proposes that habit formation is driven by action prediction errors which occur when non-habitual actions are taken. It has been further suggested that action prediction errors are encoded in activity of specific dopamine neurons, and it has been recently observed that dopamine activity in the tail of the striatum follows a pattern consistent with the action prediction errors. However, the original models capture changes in habits across trials, but do not describe the time-course of action prediction errors within trials, hence it is difficult to directly compare them with dopamine activity. We begin by outlining the 'temporal-difference action learning' algorithm, which uses biologically-plausible mechanisms to determine how dynamic changes in action intensity influence the resultant prediction errors across near-continuous time. We then demonstrate that dopaminergic data recently collected from the tail of the striatum is better represented by action prediction errors than reward prediction errors. Overall, our results support the existence of value-free action prediction errors and associated habitual behaviour in dopaminergic signals.</jats:p>