A bilingual study of socio-cultural bias in large language models through BanglaBBQ and a post processing mitigation pipeline

Citation

Abstract

Large language models have achieved impressive progress in natural language understanding, but their application in practice still brings to light a vexed and understudied issue: social bias. The majority of existing bias benchmarks were constructed with largely Western, English-centric contexts, and low-resource languages and culturally diverse societies have a big gap. This gap is filled in this paper by two related contributions. We present our first bias assessment benchmark, first, the BanglaBBQ, the first bias assessment system tailored to the Bangladeshi sociocultural environment, with nine types of bias, four of them adapted to the original BBQ framework, and five newly created, such as Regional Affiliation, Educational Background, Marital Status, Mental Health, and Politics, based on recorded sociocultural realities of Bangladesh. The dataset is bilingual with structurally aligned entries in English and Bengali allowing comparison across languages. Second, we introduce SafeLLM, a threestep inference-time bias mitigation pipeline that can be trained without retraining models or having access to weights. SafeLLM uses a sensitivity layer restructuring prompts and then inferring, a bias evaluator indicating stereotype-based predictions on the sample-level and a counterfactual grounding phase that fixes identity-sensitive mistakes by exchanging the features under protection and sampling the output. Four multilingual LLMs (LLaMA- 3.1-8B, LLaMA-3.3-70B, LLaMA-4-Scout-17B, and Gemini 2.0 Flash-Lite) are tested on both languages and all categories of bias. We find a pattern of performance decreases on culturally specific templates, significant cross-lingual accuracy differences, and a model-scale dependence in the effectiveness of inference-time interventions to decrease bias. Collectively, both BanglaBBQ and SafeLLM provide a basis to culture-specific bias measurement and mitigation in multilingual AI systems.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 69-72).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International