R-SPEAK, an AI-Based Smartphone App to Enhance Communication in People With Expressive Aphasia: Qualitative Acceptability and Usability Study

Journal of Medical Internet Research · · g524901

Background: Aphasia, an acquired language disorder that affects the ability to understand and produce language, significantly impairs effective communication. Large language models (LLMs) such as ChatGPT may help by generating fluent and coherent text, offering new ways to support communication for people with aphasia. Objective: This study aims to coproduce a communication support system using LLMs and evaluate its usefulness and acceptability among people with mild-to-moderate expressive aphasia. Methods: We used the Double Diamond approach across 3 phases. Phase 1: a stroke-survivor patient and public involvement (PPI) group (n=4) and the research team used MoSCoW prioritization (Must Have, Should Have, Could Have, Will Not Have) to rank ideas and co-design a software solution (Revolutionizing Speech Enhancement in Aphasia Using Knowledgeable-AI [R-SPEAK]) to augment verbal communication. Phase 2: 8 LLMs were evaluated on interpreting aphasic utterances from AphasiaBank transcripts. Outputs were compared, and the best responses, selected by team consensus, formed ground truth. The model producing the highest-quality responses was used in the prototype. Four people with aphasia and one carer (n=5) evaluated the prototype in semistructured interviews, and a health care professional (HCP) focus group (n=6) evaluated the concept and prototype. Topic guides were informed by the technology acceptance model (TAM), and thematic analysis themes were mapped onto its constructs. People with aphasia (n=4) rated prototype usability using the System Usability Scale (SUS). Phase 3: to improve processing speed, 12 lightweight LLMs (0.5-4 billion parameters) were evaluated on interpreting real aphasic speech using an LLM-as-a-judge framework to assess answer relevancy, faithfulness, and completeness. Results: In Phase 1, PPI groups (people living with aphasia, carers, and HCPs) identified essential system features and contextual data requirements across 2 co-ideation sessions using MoSCoW prioritization. In Phase 2, Mixtral 8x7B, the best-performing LLM for interpreting aphasic utterances, was used for the prototype. People living with aphasia rated R-SPEAK as good (SUS score: mean 75, SD 9.19). Themes mapped across three TAM constructs. (1) Attitude toward using it: people living with aphasia had high hopes, while clinicians were more cautious about its benefits. (2) Perceived ease of use: participants found it easy, though potentially harder for those with other poststroke impairments or more severe aphasia, with training possibly needed. (3) Perceived usefulness: R-SPEAK could be useful in many scenarios and improve independence. Recommendations included improved accuracy, speed, and tailored interface modifications. Phase 3 showed Qwen2.5:3B performed strongest overall, with high faithfulness and sub-second latency, while models under 1.5 billion parameters showed pronounced hallucination, indicating a lower bound on model capacity for reliable clinical speech interpretation. Conclusions: Our co-designed AI-supported R-SPEAK prototype was considered acceptable to patients. Next steps involve refining and developing a smartphone app for feasibility testing in a larger cohort of people with mild-to-moderate aphasia.

Background: Aphasia, an acquired language disorder that affects the ability to understand and produce language, significantly impairs effective communication. Large language models (LLMs) such as ChatGPT may help by generating fluent and coherent text, offering new ways to support communication for people with aphasia. Objective: This study aims to coproduce a communication support system using LLMs and evaluate its usefulness and acceptability among people with mild-to-moderate expressive aphasia. Methods: We used the Double Diamond approach across 3 phases. Phase 1: a stroke-survivor patient and public involvement (PPI) group (n=4) and the research team used MoSCoW prioritization (Must Have, Should Have, Could Have, Will Not Have) to rank ideas and co-design a software solution (Revolutionizing Speech Enhancement in Aphasia Using Knowledgeable-AI [R-SPEAK]) to augment verbal communication. Phase 2: 8 LLMs were evaluated on interpreting aphasic utterances from AphasiaBank transcripts. Outputs were compared, and the best responses, selected by team consensus, formed ground truth. The model producing the highest-quality responses was used in the prototype. Four people with aphasia and one carer (n=5) evaluated the prototype in semistructured interviews, and a health care professional (HCP) focus group (n=6) evaluated the concept and prototype. Topic guides were informed by the technology acceptance model (TAM), and thematic analysis themes were mapped onto its constructs. People with aphasia (n=4) rated prototype usability using the System Usability Scale (SUS). Phase 3: to improve processing speed, 12 lightweight LLMs (0.5-4 billion parameters) were evaluated on interpreting real aphasic speech using an LLM-as-a-judge framework to assess answer relevancy, faithfulness, and completeness. Results: In Phase 1, PPI groups (people living with aphasia, carers, and HCPs) identified essential system features and contextual data requirements across 2 co-ideation sessions using MoSCoW prioritization. In Phase 2, Mixtral 8x7B, the best-performing LLM for interpreting aphasic utterances, was used for the prototype. People living with aphasia rated R-SPEAK as good (SUS score: mean 75, SD 9.19). Themes mapped across three TAM constructs. (1) Attitude toward using it: people living with aphasia had high hopes, while clinicians were more cautious about its benefits. (2) Perceived ease of use: participants found it easy, though potentially harder for those with other poststroke impairments or more severe aphasia, with training possibly needed. (3) Perceived usefulness: R-SPEAK could be useful in many scenarios and improve independence. Recommendations included improved accuracy, speed, and tailored interface modifications. Phase 3 showed Qwen2.5:3B performed strongest overall, with high faithfulness and sub-second latency, while models under 1.5 billion parameters showed pronounced hallucination, indicating a lower bound on model capacity for reliable clinical speech interpretation. Conclusions: Our co-designed AI-supported R-SPEAK prototype was considered acceptable to patients. Next steps involve refining and developing a smartphone app for feasibility testing in a larger cohort of people with mild-to-moderate aphasia.

Источник: Journal of Medical Internet Research