Abstract
<title>Abstract</title> <p>We examine a symmetric multi-agent Bertrand pricing game to demonstrate that a Thompson Sampling (TS) learning algorithm—when coupled with a simple imitation mechanism—converges reliably to the Nash equilibrium. We prove convergence and validate the result numerically. TS finds the equilibrium in a decentralized, model-free manner, i.e., no player needs knowledge of the payoff structure or gradient information, thereby offering a practical approach to reinforcement learning in multi-agent settings with a common payoff structure and a unique equilibrium. We view this as a first step toward using TS-based learning in richer multi-agent environments where analytical Nash solutions are difficult or infeasible to derive. In that context, we point to the appeal of embedding such learning mechanisms in agent-based models. The codes to replicate the results presented here are available with this paper.</p>