Multimodal Emotion Recognition in Healthcare Dialogue Using Language and Physiological Signals
Keywords:
multimodal emotion recognition; healthcare dialogue; physiological signals; natural language processing; affective computing; clinical artificial intelligence; fairness; governanceAbstract
Emotion recognition in healthcare dialogue has emerged as a significant interdisciplinary challenge that intersects natural language processing, physiological sensing, clinical informatics, and systems engineering. Unlike conventional affective computing, clinical emotion recognition must operate under strict requirements for reliability, privacy, interpretability, and equitable performance across heterogeneous patient populations. This paper presents a system-level examination of multimodal emotion recognition that fuses language from clinical conversations with physiological signals such as heart rate variability, electrodermal activity, and skin temperature. The discussion addresses architectural choices in multimodal fusion, the trade-offs between early and late integration, and the infrastructural demands of deploying such systems in real healthcare environments. It further considers signal quality, temporal alignment, missing modality handling, and the limitations of supervised models under dataset shift. The paper argues that clinical emotion recognition cannot be treated solely as a classification problem; it must be governed as a socio-technical infrastructure that operates within regulatory frameworks, clinical workflows, and ethical boundaries. Robustness, fairness, privacy, and auditability are therefore treated as first-class design constraints rather than post hoc considerations. The work draws on cross-domain examples from digital mental health, chronic disease management, and clinical language modeling to illustrate the structural pressures that shape deployment. It concludes by outlining future directions in adaptive fusion, continual learning, and governance-oriented evaluation.
References
1. Poria, S., Cambria, E., Bajpai, R., & Hussain, A. (2017). A review of affective computing: From unimodal analysis to multimodal fusion. Information Fusion, 37, 98–125. https://doi.org/10.1016/j.inffus.2017.02.003
2. Picard, R. W. (1997). Affective computing. MIT Press.
3. Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7
4. Hirschberg, J., & Manning, C. D. (2015). Advances in natural language processing. Science, 349(6245), 261–266. https://doi.org/10.1126/science.aaa8685
5. Li, Q. (2026). Dynamic Adaptive Attention and Supervised Contrastive Learning: A Novel Hybrid Framework for Text Sentiment Classification. arXiv preprint arXiv:2604.10459.
6. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
7. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–4186.
8. Subramanian, R., Wache, J., Abadi, M. K., Vieriu, R. L., Winkler, S., & Sebe, N. (2018). ASCERTAIN: Emotion and personality recognition using commercial physiological sensors. IEEE Transactions on Affective Computing, 9(2), 147–160. https://doi.org/10.1109/TAFFC.2016.2625250
9. Schuller, B., & Batliner, A. (2013). Computational paralinguistics: Emotion, affect and personality in speech and language processing. Wiley.
10. Busso, C., Deng, Z., Yildirim, S., Bulut, M., Lee, C. M., Kazemzadeh, A., Lee, S., Neumann, U., & Narayanan, S. (2004). Analysis of emotion recognition using facial expressions, speech and multimodal information. Proceedings of the 6th International Conference on Multimodal Interfaces, 205–211. https://doi.org/10.1145/1027933.1027968
11. Kächele, M., Schels, M., & Schwenker, F. (2014). Adaptive confidence fusion of heterogeneous classifiers for emotion recognition. Pattern Recognition Letters, 46, 1–7. https://doi.org/10.1016/j.patrec.2014.04.011
12. Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x
13. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607
14. Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G., & Chin, M. H. (2018). Ensuring fairness in machine learning to advance health equity. Annals of Internal Medicine, 169(12), 866–872. https://doi.org/10.7326/M18-1990
15. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
16. Voigt, P., & Von dem Bussche, A. (2017). The EU General Data Protection Regulation (GDPR): A practical guide. Springer.
17. United States Department of Health and Human Services. (2003). Summary of the HIPAA Privacy Rule. Office for Civil Rights.
18. Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2), 1–21. https://doi.org/10.1177/2053951716679679
19. Miotto, R., Wang, F., Wang, S., Jiang, X., & Dudley, J. T. (2018). Deep learning for healthcare: Review, opportunities and challenges. Briefings in Bioinformatics, 19(6), 1236–1246. https://doi.org/10.1093/bib/bbx044
20. Doryab, A., Frost, M., Faurholt-Jepsen, M., Kessing, L. V., & Bardram, J. E. (2015). Impact factor analysis of multi-modal data for monitoring bipolar disorder. IEEE Pervasive Computing, 14(2), 30–38. https://doi.org/10.1109/MPRV.2015.50
21. Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., & Wermter, S. (2019). Continual lifelong learning with neural networks: A review. Neural Networks, 113, 54–71. https://doi.org/10.1016/j.neunet.2019.01.012
22. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.