← Projects
Fairness in LLM Toxicity Detection
Evaluated attributional bias in LLM toxicity detection using counterfactual author identities, measuring attribution shift, false-positive disparity, alignment bias, calibration, and differences between GPT-OSS-120B and LLaMA-3.1-8B.