Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>This paper investigates the detection and annotation of bias in Bulgarian with a view to large language models (LLMs). We present a Bias Annotated Dataset in Bulgarian comprising 3,177 sentences extracted from Bulgarian Wikipedia, manually annotated by two native speakers on five bias types (gender, religion, race and ethnicity, physical appearance, disability, and an additional ’Other’ category) using an ordinal scale from 0 to 5. We evaluate inter-annotator agreement both between the two human annotators and between humans and LLMs (Gemma 4, BgGPT-Gemma-3 with English and Bulgarian prompting, and Qwen 3) using weighted Cohen’s κ, Krippendorff’s α, and intraclass correlation coefficient (ICC). The results reveal high agreement between human annotators on the binary detection task (biased / non-biased classification) with mismatches of only 4.9%, but substantially lower agreement on intensity scoring, consistent with the inherently subjective nature of bias perception. LLMs show significant divergence from human judgements, with much higher mismatch rates of 24–34% and uniformly negative agreement metrics, indicating that LLMs flag largely different sentences as biased and assign systematically different (lower) intensity scores than human annotators. With respect to the BgGPT- Gemma-3 models specifically focused on Bulgarian, the results show that prompting in Bulgarian rather than English improves LLM alignment with human annotation, highlighting the importance of target-language prompting for low-resource and culturally specific bias detection. The findings contribute to the understanding of bias in Bulgarian and low-resource languages more broadly, and underscore the need for culturally grounded lexical resources and annotation methodologies.</p>

Show More

Keywords

bulgarian bias human llms agreement

Related Articles

PORE

About

Connect