Fairness-Aware Prompt Tuning: Selective Prompt Allocation for Bias Reduction in NLP Models
Keywords:
fairness-aware NLP, prompt tuning, selective prompt allocation, bias mitigation, socio-technical systems, language model deploymentAbstract
Prompt tuning has emerged as a parameter-efficient alternative to full model fine-tuning, enabling large language models to be adapted to downstream tasks through learned soft prompts while keeping pretrained weights frozen. Despite its efficiency advantages, the fairness implications of prompt tuning remain underexplored. Standard prompt tuning applies a single prompt uniformly across all instances, implicitly assuming that a monolithic prompt representation can serve all demographic subgroups equally well. This assumption is at odds with extensive evidence that language models amplify societal biases embedded in training data. In this paper, we introduce the concept of fairness-aware prompt tuning through selective prompt allocation, a system-level framework in which inputs are routed to distinct fairness-optimized prompts based on auxiliary signals that capture protected attribute proxies or bias-sensitive features. We analyze the structural trade-offs involved in partitioning the prompt space for fairness, including prompt catalog design, allocation policy learning, and the tension between group fairness and aggregate accuracy. We further examine infrastructure requirements for deploying selective prompt allocators in production pipelines, touching on latency, storage overhead, and the risks of re-identification when conditioning on sensitive attributes. The governance dimension is discussed in terms of regulatory compliance, transparency obligations, and the necessity of auditability in prompt-based adaptation layers. Through cross-domain comparisons with adversarial debiasing, data augmentation, and output calibration methods, we illustrate how selective prompt allocation creates a new operating point in the bias-accuracy landscape. We conclude by identifying challenges related to long-term fairness drift, adversarial prompt manipulation, and sustainable maintenance, suggesting that selective prompt allocation represents a promising middleware approach for fairness injection in large-scale NLP systems without compromising the efficiency gains of prompt tuning.
References
1. Lester, B., Al-Rfou, R., & Constant, N. (2021). The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 3045–3059).
2. Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2022). GPT understands, too. AI Open, 3, 62–70.
3. Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (pp. 4582–4597).
4. Bolukbasi, T., Chang, K. W., Zou, J., Saligrama, V., & Kalai, A. (2016). Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Advances in Neural Information Processing Systems (Vol. 29, pp. 4349–4357).
5. Zhao, J., Wang, T., Yatskar, M., Ordonez, V., & Chang, K. W. (2017). Men also like shopping: Reducing gender bias amplification using corpus-level constraints. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (pp. 2979–2989).
6. Sun, T., Gaut, A., Tang, S., Huang, Y., ElSherief, M., Zhao, J., Mirza, D., Belding, E., Chang, K. W., & Wang, W. Y. (2019). Mitigating gender bias in natural language processing: Literature review. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 1630–1640).
7. Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems (Vol. 29, pp. 3315–3323).
8. Dwork, C., Hardt, M., Pitassi, T., Reingold, O., & Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (pp. 214–226).
9. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), Article 115.
10. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).
11. Barocas, S., Hardt, M., & Narayanan, A. (2019). Fairness and machine learning: Limitations and opportunities. fairmlbook.org.
12. Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., & Smith, N. A. (2020). Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305.
13. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901).
14. Schick, T., & Schütze, H. (2021). Exploiting cloze questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp. 255–269).
15. Zhu, W., & Tan, M. (2023, December). SPT: Learning to selectively insert prompts for better prompt tuning. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 11862-11878).
16. Liang, P. P., Wu, C., Morency, L. P., & Salakhutdinov, R. (2021). Towards understanding and mitigating social biases in language models. In Proceedings of the 38th International Conference on Machine Learning (pp. 6565–6576).
17. Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., & Smith, N. A. (2018). Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 107–112).
18. Madan, S., Henry, T., Dozier, J., Ho, H., Bhandari, N., Sasaki, T., Durand, F., Pfister, H., & Boix, X. (2022). On the fairness of disentangled representations. In Advances in Neural Information Processing Systems (Vol. 35, pp. 30478–30491).
19. Sheng, E., Chang, K. W., Natarajan, P., & Peng, N. (2019). The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 3407–3412).
20. Tamkin, A., Brundage, M., Clark, J., & Ganguli, D. (2021). Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503.
21. Veale, M., & Binns, R. (2017). Fairer machine learning in the real world: Mitigating discrimination without collecting sensitive data. Big Data & Society, 4(2), 1–17.
22. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59–68).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.