Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Back to Search View Original Cite This Article

Abstract

<p>Deploying large language models (LLMs) in ethically sensitive roles - therapeutic support, clinical decision assistance, and legal analysis, raises the question of whether these systems possess genuine moral wisdom or merely simulate it. This study introduces a dual-layer computational psychometrics pipeline that combines lexical-emotional profiling with a five-dimension, behaviorally anchored wisdom rubric (Cognitive Integration, Contextual Adaptability, Reflective Empathic Action, Social-Emotional Equilibrium, and Ethical Flexibility) and validates it against an independent human expert baseline before using it to benchmark artificial intelligence. Free-text responses from 100 participants (1,300 vignette-level observations) to thirteen pictorial moral dilemmas were scored by five trained human raters, establishing a human-human inter-rater reliability benchmark (ICC = .764) prior to any AI comparison. Three architecturally distinct LLMs were then validated as automated raters against this human baseline, and a self-serving-bias check confirmed rater impartiality. Five state-of-the-art LLMs (GPT-4o, GPT-4o-mini, Llama-3.3-70B, Claude Sonnet, and Hy3) were subsequently benchmarked against the validated human baseline. All five systematically overrated their own moral reasoning relative to humans across every wisdom dimension, an alignment artifact further characterized by a cross-vignette variance paradox and scenario-insensitive moral differentiation consistent with reinforcement-learning-from-human-feedback-style training; a sixth model, the small language model Llama-3.1-8B, was subsequently evaluated post hoc to test whether this pattern depended on model scale and showed the identical inflation pattern despite being an order of magnitude smaller. Cross-layer analyses linked this artifact to a structurally asymmetric human affective architecture in which disgust suppresses, and anticipation and trust amplify, wisdom-relevant cognition. These findings suggest that LLM moral response distributions differ systematically from validated human moral cognition baselines across multiple independent measurement layers. Critically, automated raters exhibited high internal consistency and no self-serving bias toward AI-generated text yet achieved only Poor agreement with independent human experts (ICC = .505 and .551), while agreeing more with each other (ICC = .585), a finding with direct implications for LLM-as-judge evaluation frameworks across AI ethics, psychometric assessment, and behavioral research. These results carry implications for AI ethics governance and the responsible deployment of LLMs in emotionally consequential contexts.</p>

Show More

Keywords

human moral llms wisdom against

Related Articles


Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 76
PORE

About

Connect