Research: Policy regularization as a unifying theory of the striatal division of labor in learning
Learning novel behaviors requires balancing previously learned actions with the ability to flexibly adapt to changing reward contingencies. This trade-off is well documented in the division of labor between dorsolateral striatum (DLS), which promotes selection of cached, history-dependent actions, and dorsomedial striatum (DMS), which supports flexible learning as reward contingencies change. Here, we propose policy regularization as a general computational principle for understanding this division. Capacity-limited agents face a fundamental trade-off between maximizing reward and minimizing the cost of deviating from a default policy that caches frequently used action transitions. We formalize this trade-off as a KL-regularized reward objective in which a flexible controller (DMS) incurs a cost for diverging from a history-dependent default (DLS). The resulting optimal policy is a combination of a reward-driven action value, continuously updated by DMS, and a default policy conditioned on action history, cached by DLS. In practice, DLS consolidates the action transitions shaped by DMS's reward-driven value learning, so DLS preserves behaviors that were once optimal even when reward contingencies change, producing robust but inflexible action selection and effectively "regularizing" DMS-driven learning. We show that a single model with one shared parameter set and consistent lesion rules provides a unifying explanation for the functional organization of the striatum, reproducing canonical DLS-DMS dissociations in outcome devaluation, serial spatial reversal, skilled action sequencing, and motor sequence execution tasks.