Barcelona, Veronica; Scharp, Danielle; Moen, Hans; Davoudi, Anahita; Idnay, Betina R.; Cato, Kenrick; Topaz, Maxim

doi:10.1007/s10995-023-03857-4

Back to matches

Your institution may have access to this item. Find your institution then sign in to continue.

Title: Using Natural Language Processing to Identify Stigmatizing Language in Labor and Birth Clinical Notes.
Authors: Barcelona, Veronica; Scharp, Danielle; Moen, Hans; Davoudi, Anahita; Idnay, Betina R.; Cato, Kenrick; Topaz, Maxim
Abstract: Introduction: Stigma and bias related to race and other minoritized statuses may underlie disparities in pregnancy and birth outcomes. One emerging method to identify bias is the study of stigmatizing language in the electronic health record. The objective of our study was to develop automated natural language processing (NLP) methods to identify two types of stigmatizing language: marginalizing language and its complement, power/privilege language, accurately and automatically in labor and birth notes. Methods: We analyzed notes for all birthing people > 20 weeks' gestation admitted for labor and birth at two hospitals during 2017. We then employed text preprocessing techniques, specifically using TF-IDF values as inputs, and tested machine learning classification algorithms to identify stigmatizing and power/privilege language in clinical notes. The algorithms assessed included Decision Trees, Random Forest, and Support Vector Machines. Additionally, we applied a feature importance evaluation method (InfoGain) to discern words that are highly correlated with these language categories. Results: For marginalizing language, Decision Trees yielded the best classification with an F-score of 0.73. For power/privilege language, Support Vector Machines performed optimally, achieving an F-score of 0.91. These results demonstrate the effectiveness of the selected machine learning methods in classifying language categories in clinical notes. Conclusion: We identified well-performing machine learning methods to automatically detect stigmatizing language in clinical notes. To our knowledge, this is the first study to use NLP performance metrics to evaluate the performance of machine learning methods in discerning stigmatizing language. Future studies should delve deeper into refining and evaluating NLP methods, incorporating the latest algorithms rooted in deep learning. Significance: What is Already Known on this Subject?: Traditional informatics methods include natural language processing, and these methods have been increasingly applied to the study of public health problems using electronic health records. What this Study Adds?: We identified well-performing machine learning methods to automatically identify stigmatizing language in labor and birth clinical notes. These methods have not been applied to labor and birth clinical notes and have the potential to be a powerful tool in examining perinatal health inequities.
Subjects: MATERNAL health services; PATIENT abuse; NATURAL language processing; RESEARCH methodology; MACHINE learning; SOCIAL stigma; LANGUAGE &; languages; ACQUISITION of data; QUALITATIVE research; MEDICAL records; RESEARCH funding; ELECTRONIC health records; THEMATIC analysis; DATA mining; SECONDARY analysis; ALGORITHMS
Publication: Maternal & Child Health Journal, 2024, Vol 28, Issue 3, p578
ISSN: 1092-7875
Publication type: Article
DOI: 10.1007/s10995-023-03857-4

We found a match

Using Natural Language Processing to Identify Stigmatizing Language in Labor and Birth Clinical Notes.

Barcelona, Veronica; Scharp, Danielle; Moen, Hans; Davoudi, Anahita; Idnay, Betina R.; Cato, Kenrick; Topaz, Maxim

Maternal & Child Health Journal, 2024, Vol 28, Issue 3, p578

1092-7875

Article

10.1007/s10995-023-03857-4