Back to Search View Original Cite This Article

Abstract

<p>Artificial intelligence systems are increasingly evaluated as behavioral agents, yet many evaluation practices still treat model outputs as isolated responses to be scored for accuracy, safety, bias, or preference alignment. We argue that this response-level paradigm is insufficient for systems whose behavior is generated from probability distributions and changes across contexts, versions, and interventions. Human psychometrics and cognitive modeling offer a useful foundation because they were developed to infer latent traits, capacities, preferences, and decision processes from indirect behavioral data. However, artificial agents create a different measurement regime. Their next-token and response-level probability distributions can often be inspected directly, allowing uncertainty, calibration, bias, and response instability to be modeled as latent properties rather than inferred only from overt responses. Artificial agents can also be remeasured and intervened on under controlled conditions, enabling parameter-level assessment of stability, drift, improvement, degradation, and causal sensitivity across prompts, model updates, fine-tuning, retrieval changes, and ablations. We propose machine psychometrics as a distribution-aware measurement framework for artificial agents. In this framework, computational cognitive models serve as interpretable interfaces between probability distributions and latent parameters, allowing model behavior to be compared, validated, and monitored without assuming that human psychological constructs transfer unchanged to machines. This approach shifts AI evaluation from scoring what a model says to measuring the probabilistic structure that generates its behavior.</p>

Show More

Keywords

artificial agents model from behavior

Related Articles

PORE

About

Connect