Fairness in LLM Toxicity Detection
Oct 2025 - Dec 2025Evaluated attributional bias in LLM toxicity detection using counterfactual author identities, measuring attribution shift, false-positive disparity, alignment bias, calibration, and differences between GPT-OSS-120B and LLaMA-3.1-8B.
- Python
- LLM Evaluation