Efficiency vs Depth in AI-Generated and Human-Synthesized Thematic Analyses of Ugandan Women’s Experiences of Obstetric Fistula: Comparative Qualitative Study

Journal of Medical Internet Research ·

Background: Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations. Objective: Our research sought to (1) describe the process of using AI to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research. Methods: Nested within a larger mixed methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semistructured interviews exploring their mental, physical, and overall health trajectories. Each interview lasted about an hour and was audio-recorded. Following translation and transcription, the data underwent thematic analysis by humans and by Versa, a University of California, San Francisco (UCSF) developed LLM powered by ChatGPT-4o. The AI analysis was conducted using 2 strategies: inductively and deductively. Finally, the analytic outputs were compared by MP and reviewed by AME and MG to evaluate the code frequency and alignment (most frequently used codes and equivalent concepts vs unique concepts between human and AI codes), thematic robustness (differentiation of distinct concepts and labels that convey substantive findings), analysis quality (narrative depth and excerpt accuracy), and efficiency of each method (total person-hours spent on comparable analytic tasks). Results: A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged. When inductively coding, Versa identified 39 codes, whereas human researchers used a more expansive set of 54 codes. While Versa’s thematic analyses were generally accurate regarding how the interview data were represented in the analyses, the human-synthesized analysis was more robust because it incorporated compelling excerpts, and the summary content provided greater depth and range that the Versa analyses lacked. The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.5 hours. Conclusions: While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat. Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences. Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed to maintain ethical and interpretive rigor. Trial Registration: ClinicalTrials.gov NCT05437939; https://clinicaltrials.gov/study/NCT05437939

Background: Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations. Objective: Our research sought to (1) describe the process of using AI to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research. Methods: Nested within a larger mixed methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semistructured interviews exploring their mental, physical, and overall health trajectories. Each interview lasted about an hour and was audio-recorded. Following translation and transcription, the data underwent thematic analysis by humans and by Versa, a University of California, San Francisco (UCSF) developed LLM powered by ChatGPT-4o. The AI analysis was conducted using 2 strategies: inductively and deductively. Finally, the analytic outputs were compared by MP and reviewed by AME and MG to evaluate the code frequency and alignment (most frequently used codes and equivalent concepts vs unique concepts between human and AI codes), thematic robustness (differentiation of distinct concepts and labels that convey substantive findings), analysis quality (narrative depth and excerpt accuracy), and efficiency of each method (total person-hours spent on comparable analytic tasks). Results: A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged. When inductively coding, Versa identified 39 codes, whereas human researchers used a more expansive set of 54 codes. While Versa’s thematic analyses were generally accurate regarding how the interview data were represented in the analyses, the human-synthesized analysis was more robust because it incorporated compelling excerpts, and the summary content provided greater depth and range that the Versa analyses lacked. The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.5 hours. Conclusions: While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat. Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences. Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed to maintain ethical and interpretive rigor. Trial Registration: ClinicalTrials.gov NCT05437939; https://clinicaltrials.gov/study/NCT05437939

Источник: Journal of Medical Internet Research