Abstract
<title>Abstract</title> <p>AI research agents, which are claimed to accelerate or even replace human bioinformatics analysis, are being deployed rapidly in life sciences and biomedical research. However, many scientists lack a systematic approach to assessing their utility and to determining which agent is best suited to a given task. Here we quantitatively benchmark three biomedical AI research agents, including K-Dense (Gemini 2.5 Pro), Finch from Edison Scientific, and Biomni (Claude family), against ChatGPT, which our survey identified as the most commonly used approach for analysing omics datasets. For benchmarking, we performed bioinformatic analyses of metabolomic, bulk proteomic and bulk transcriptomic datasets, which are among the most frequently requested workflows submitted to the deployed AI agents. We evaluated all four agents using a process-anchored framework that scores each output on process quality, result reporting, and interpretation depth, while a domain-aware ordinal rubric is applied in parallel. Eight independent runs per agent per modality, together with bootstrap resampling, alternative scoring definitions, output-volume controls, and within-agent reproducibility measurements, quantified uncertainty. No single agent dominated across all modalities. K-Dense led in metabolomics with the most thorough, reproducible reports; Finch produced the best-organized reports and won proteomics and transcriptomics; Biomni offered the strongest biological interpretation but frequently omitted quality control and preprocessing; ChatGPT, without workflow scaffolding, managed only basic statistics and finished last.Furthermore, agent architecture, not base-model capability, dominates what each agent outputs, with process-anchored evaluation reordering the field relative to accuracy alone. Our benchmark and rubric provide the community with an auditable basis for choosing among AI research agents for omics analysis.</p>