Back to Search View Original Cite This Article

Abstract

<p>This exploratory pilot study examined performance under standardized AI recommendations in brief workplace- style software-maintenance tasks. Thirteen adults in technology- related roles were assigned nonrandomly, using a predetermined intake-order sequence, to a no-recommendation control (n = 4), a standardized-recommendation condition (n = 4), or the same recommendations paired with a metacognitive scaffold (n = 5). Participants completed three procedural tasks, three judgment tasks, and two immediate unaided near-transfer tasks. Recommendations were fixed and researcher-prepared rather than generated through live interaction with a large language model. Descriptively, the unscaffolded recommendation group completed the main tasks fastest (71.5 s) and exceeded the control on procedural accuracy (11/12 versus 8/12 correct), but showed lower judgment accuracy (2/12 correct), greater signed overconfidence (+33.3 percentage points), and lower immediate transfer (2/8 correct). The scaffold group was slower (118.4 s) but showed higher judgment accuracy (12/15 correct), better confidence calibration (mean Brier score 0.058), and higher immediate transfer (9/10 correct). These results are feasibility evidence rather than causal estimates of typical AI use. The sample was small, assignment was nonrandom, task scales were short, recommendation correctness was confounded with task category, and the scaffold combined additional time, written elaboration, and explicit error search. The observed pattern motivates larger studies that independently manipulate recommendation correctness and verification support.</p>

Show More

Keywords

tasks correct recommendations scaffold judgment

Related Articles

PORE

About

Connect