Suicide-Related Responses From AI Chatbots Through Application Programming Interfaces and Consumer-Facing Apps: Observational Cross-Sectional Comparative Study

Journal of Medical Internet Research ·

Background: Generative artificial intelligence (GAI) chatbots are frequently used for mental health–related questions, including suicide-related queries. User interfaces (UIs) may include additional safety controls not present in direct application programming interface (API) access (eg, moderation rules that flag or block harmful content, crisis refusal templates, or system-level instructions). Objective: This observational cross-sectional study aimed to examine how frequently AI chatbots provide direct responses to suicide-related prompts across access modes and whether response patterns differ between consumer-facing UI and API access across models and prompt risk levels. Methods: We evaluated how five widely used consumer AI models (ChatGPT, Claude, Gemini, Grok, and Llama) responded to 30 previously vetted suicide-related prompts spanning five clinician-assigned risk levels. Each prompt was submitted 100 times through both public-facing UI and direct API access, yielding 30,000 total responses. The primary outcome was whether a response directly answered the prompt (ie, provided specific information or guidance related to the question asked). AI responses were categorized using a blinded large language model–based classifier (GPT-4o mini), validated against independent human coding (289/300, 96.3% agreement; Cohen κ=0.89). We estimated mixed-effects logistic regression models that predicted a direct response. The primary predictors were the AI model, access mode (UI vs API), and prompt risk category. Results: Of the 30,000 responses, 69.8% (n=20,934) were direct. Direct responses were more common through APIs than UIs (n=11,609, 77.4%, vs n=9325, 62.2%); however, Gemini had slightly higher direct response rates via UIs than APIs (n=2301, 76.7%, vs n=2131, 71.0%). Differences in the likelihood of a direct response were most pronounced for higher-risk prompts: For very-high-risk prompts, 24.8% (n=743) of API responses were direct compared with 4.6% (n=138) of UI responses, whereas for high-risk prompts, 80.1% (n=2002) of API responses were direct compared with 48.4% (n=1209) of UI responses. Claude and Gemini had the highest direct response rates (n=4688, 78.1%, and n=4432, 73.9%, respectively) across UI and API access modes, whereas ChatGPT, Grok, and Llama had lower rates (n=3878, 64.6%; n=3870, 64.5%; and n=4066, 67.8%, respectively). In mixed-effects models, UI access was associated with lower odds of a direct response than API access (odds ratio 0.09, 95% CI 0.08-0.10). Higher prompt risk was associated with lower direct response probability, and access mode differences varied by risk level and model. Conclusions: AI safety depends on how users access the model. Safety observed in a UI should not be assumed to generalize to API access or apps built on APIs. Evaluations and policies should consider access modes to ensure safety protections for individuals interacting with AI chatbots.

Background: Generative artificial intelligence (GAI) chatbots are frequently used for mental health–related questions, including suicide-related queries. User interfaces (UIs) may include additional safety controls not present in direct application programming interface (API) access (eg, moderation rules that flag or block harmful content, crisis refusal templates, or system-level instructions). Objective: This observational cross-sectional study aimed to examine how frequently AI chatbots provide direct responses to suicide-related prompts across access modes and whether response patterns differ between consumer-facing UI and API access across models and prompt risk levels. Methods: We evaluated how five widely used consumer AI models (ChatGPT, Claude, Gemini, Grok, and Llama) responded to 30 previously vetted suicide-related prompts spanning five clinician-assigned risk levels. Each prompt was submitted 100 times through both public-facing UI and direct API access, yielding 30,000 total responses. The primary outcome was whether a response directly answered the prompt (ie, provided specific information or guidance related to the question asked). AI responses were categorized using a blinded large language model–based classifier (GPT-4o mini), validated against independent human coding (289/300, 96.3% agreement; Cohen κ=0.89). We estimated mixed-effects logistic regression models that predicted a direct response. The primary predictors were the AI model, access mode (UI vs API), and prompt risk category. Results: Of the 30,000 responses, 69.8% (n=20,934) were direct. Direct responses were more common through APIs than UIs (n=11,609, 77.4%, vs n=9325, 62.2%); however, Gemini had slightly higher direct response rates via UIs than APIs (n=2301, 76.7%, vs n=2131, 71.0%). Differences in the likelihood of a direct response were most pronounced for higher-risk prompts: For very-high-risk prompts, 24.8% (n=743) of API responses were direct compared with 4.6% (n=138) of UI responses, whereas for high-risk prompts, 80.1% (n=2002) of API responses were direct compared with 48.4% (n=1209) of UI responses. Claude and Gemini had the highest direct response rates (n=4688, 78.1%, and n=4432, 73.9%, respectively) across UI and API access modes, whereas ChatGPT, Grok, and Llama had lower rates (n=3878, 64.6%; n=3870, 64.5%; and n=4066, 67.8%, respectively). In mixed-effects models, UI access was associated with lower odds of a direct response than API access (odds ratio 0.09, 95% CI 0.08-0.10). Higher prompt risk was associated with lower direct response probability, and access mode differences varied by risk level and model. Conclusions: AI safety depends on how users access the model. Safety observed in a UI should not be assumed to generalize to API access or apps built on APIs. Evaluations and policies should consider access modes to ensure safety protections for individuals interacting with AI chatbots.

Источник: Journal of Medical Internet Research