Fairness-Aware Benchmarking of Cross-Cultural Visual Generation in Foundation Models
Keywords:
foundation models, text-to-image generation, fairness benchmarking, cross-cultural representation, sociotechnical systems, visual bias, AI governanceAbstract
As foundation models increasingly mediate visual culture across global audiences, their capacity to fairly represent diverse cultural contexts has emerged as a pressing socio-technical challenge. Current text-to-image generation systems exhibit profound demographic and cultural biases, often defaulting to Western, high-resource visual norms while erasing or stereotyping underrepresented communities. This paper presents a system-level analysis of fairness-aware benchmarking for cross-cultural visual generation, arguing that effective evaluation requires far more than dataset curation or prompt engineering. We articulate how benchmark design must integrate infrastructure scalability, metric pluralism, governance scaffolding, and sustainability considerations into a coherent sociotechnical architecture. Through a critical examination of existing evaluation pipelines, we identify structural trade-offs between cultural granularity and computational tractability, between automated metrics and culturally grounded human evaluation, and between centralised model auditing and community-driven oversight. Drawing on insights from recent model deployments, we propose a layered benchmarking architecture that couples large-scale automated analysis with distributed cultural annotation networks, enabling continuous monitoring of representation gaps. We further discuss the role of regulatory frameworks, model cards, and participatory design in shaping accountability for cross-cultural fairness. Throughout, we emphasise that cross-cultural fairness in generative AI is not a static property to be optimised but a dynamic, contested terrain that demands ongoing institutional commitment, reflexive evaluation practices, and infrastructure that supports pluralism rather than erasing it. The paper contributes a comprehensive framework for thinking about fairness-aware benchmarking as a system of measurement, governance, and cultural negotiation.
References
1. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).
2. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning.
3. Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125.
4. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684–10695).
5. Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., & Caliskan, A. (2023). Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (pp. 1493–1504).
6. Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., ... Gabriel, I. (2021). Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
7. Birhane, A., Prabhu, V. U., & Kahembwe, E. (2021). Multimodal datasets: Misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963.
8. Cho, J., Zala, A., & Bansal, M. (2023). DALL-Eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 22618–22629).
9. Hershcovich, D., Frank, S., Lent, H., de Lhoneux, M., Abdou, M., Zhao, J., Yanaka, H., Bugliarello, E., Stengel-Eskin, E., Linzen, T., Tiedemann, J., Søgaard, A., & Sennrich, R. (2022). Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (pp. 6997–7013).
10. Bhatt, S., Dev, S., Talukdar, P., Dave, S., Prabhakaran, V., & Garg, N. (2022). The geography of bias in language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (pp. 1538–1550).
11. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
12. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229).
13. Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587–604.
14. Mohamed, S., Png, M.-T., & Isaac, W. (2023). Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence. Philosophy & Technology, 36(3), Article 63.
15. Prabhakaran, V., Mitchell, M., Gebru, T., & Gabriel, I. (2023). Sociotechnical perspectives on evaluating foundation models. In R. Bommasani et al. (Eds.), Perspectives on foundation models (pp. 187–214). Center for Research on Foundation Models, Stanford University.
16. D’Amour, A., Heller, K., Moldoveanu, M., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., Hormozdiari, F., Houlsby, N., Hou, S., Jerfel, G., Karthikesalingam, A., Lucic, M., Ma, Y., McLean, C., Mincu, D., ... Sculley, D. (2020). Underspecification presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395.
17. C. Shi, S. Li, S. Guo, S. Xie, W. Wu, J. Dou, C. Wu, C. Xiao, C. Wang, Z. Cheng, et al. (2025)Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282.
18. Raji, I. D., & Buolamwini, J. (2019). Actionable auditing: Investigating the impact of publicly naming biased performance results of commercial AI products. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (pp. 429–435).
19. Anthony, L. F. W., Kanding, B., & Selvan, R. (2020). Carbontracker: Tracking and predicting the carbon footprint of training deep learning models. arXiv preprint arXiv:2007.03051.
20. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.
21. Veale, M., & Borgesius, F. Z. (2021). Demystifying the draft EU Artificial Intelligence Act. Computer Law Review International, 22(4), 97–112.
22. Diakopoulos, N. (2016). Accountability in algorithmic decision making. Communications of the ACM, 59(2), 56–62.
23. Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 33–44).
24. Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., & Song, D. (2021). Natural adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 15262–15271).
25. Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., & Gilmer, J. (2021). The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 8340–8349).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.