Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<jats:p>National reports and related biodiversity policy documents contain important information on implementation barriers, but extracting comparable evidence across many countries is timeconsuming. We tested whether large language models (LLMs) could provide a reliable firstpass synthesis of reported challenges relevant to Target 6 of the Kunming–Montreal Global Biodiversity Framework. Human assessor scores for 50 countries were compared with structured outputs from three LLMs: Claude, ChatGPT and Gemini. The aim was not to test exact score reproduction, but to determine whether LLMs captured the same relative patterns in reported challenge categories. Claude showed the strongest alignment with human assessment, especially at the level of challenge-category ranking. Claude-derived scores were therefore used to summarise average challenge patterns across 126 reports. The results suggest that LLM scoring is useful for identifying broad thematic patterns in reporting challenges, but should not be used for precise country-level ranking.</jats:p>