← Projects

Fairness in LLM Toxicity Detection

Oct 2025 - Dec 2025
Evaluated attributional bias in LLM toxicity detection using counterfactual author identities, measuring attribution shift, false-positive disparity, alignment bias, calibration, and differences between GPT-OSS-120B and LLaMA-3.1-8B.