Abstract
<jats:p>We study whether content-level moderation that reduces exposure to social media toxicity, holding platform access constant, affects mental health. We deploy, for randomly assigned users, a cross-platform browser extension that hides posts and comments exceeding a prespecified toxicity threshold, and follow N = 664 desktop users over a six-week intervention across two waves – one during the 2024 U.S. election campaign and one outside of it. Contrary to basic intuition, filtering toxicity worsens mental health: treatment raises a standardized index of GAD-7 and PHQ-8 symptoms by 0.15σ, and the share of respondents above a moderate-symptom screening cutoff rises by about 12 pp. A QALY-based calibration implies a welfare cost of about $158 per treated participant. Heightened loneliness and a diminished relative moral self-view emerge as potential channels. We conclude that curbing toxic content can impose unintended psychological costs, even where it yields other social benefits.</jats:p>