
Loading, please wait...

Loading, please wait...

Cancer survivors frequently navigate complicated long-term sequelae after completing curative or maintenance oncologic treatment. Consequently, many patients turn to digital artificial intelligence platforms when they seek answers to cancer survivorship questions. These individuals often search for guidance on recurrence surveillance, chronic fatigue, medication interactions, and psychosocial adjustment. Commercial large language models now offer immediate conversational responses to these distressing queries. However, unvetted digital advice introduces serious clinical vulnerabilities, including inaccurate staging advice and inappropriate self-management instructions. In multilingual healthcare environments like Hong Kong and India, patients communicate across diverse linguistic registers and cultural idioms. Therefore, healthcare providers must understand how different conversational models handle nuanced post-treatment concerns. A landmark evaluation assessed four commercial foundation models across English and Traditional Chinese inputs. Specifically, investigators analyzed Google Gemini 3.1 Pro, HK Chat-0.6.2, DeepSeek-V3.2, and Kimi-K2.5. An expert Delphi panel evaluated twenty-five curated clinical scenarios spanning five vital survivorship domains. Ultimately, the investigators demonstrated that off-the-shelf chatbots produce divergent advice depending directly on the language queried. Clinicians must recognize these automated limitations to safeguard vulnerable oncology patients against digital misinformation.
The study thoroughly appraised whether modern artificial intelligence engines deliver practical guidance for post-therapy concerns. Furthermore, an expert Delphi panel rated baseline answers using a validated five-point clinical utility scale. Across all domains, Google Gemini 3.1 Pro achieved the highest clinical utility scores in both evaluated languages. In English interactions, the model secured an average utility score of 3.99 out of 5. Similarly, in Traditional Chinese interactions, it achieved an impressive average score of 4.21 out of 5. In contrast, regional and alternative models demonstrated inconsistent performance across complex oncology topics. DeepSeek-V3.2 and Kimi-K2.5 frequently provided superficial responses when answering queries regarding lymphedema and neuropathy management. Moreover, the investigators measured patient understandability using the Patient Education Materials Assessment Tool. Interestingly, models that scored well in structural readability did not necessarily provide superior clinical utility. For instance, highly readable outputs frequently simplified critical concepts to the point of clinical inaccuracy. Consequently, high readability metrics can disguise incomplete or misleading oncology advice. Clinicians must recognize that fluent conversational delivery does not equal therapeutic safety or technical accuracy. Therefore, medical teams must exercise caution before endorsing general artificial intelligence search tools.
Unconstrained foundation models revealed alarming safety risks that varied substantially according to the language used. Overall, ten percent of all evaluated chatbot outputs contained clinically unsafe medical guidance. However, qualitative analysis demonstrated distinct behavioral variations between English and Traditional Chinese prompts. In English responses, up to twenty-four percent of errors involved healthcare service mismatches and definitive diagnostic declarations. For example, chatbots frequently assigned absolute medical diagnoses to vague non-specific physical symptoms. Additionally, they often directed patients to inappropriate secondary services instead of recommending urgent oncology review. Conversely, Traditional Chinese responses showed a twenty percent unsafe rate characterized by prescriptive treatment instructions. Specifically, these Chinese outputs frequently promoted unverified adjunctive therapies and unvalidated herbal interventions. Such recommendations introduce grave perils of pharmacologic interactions with active adjuvant endocrine or targeted drugs. Furthermore, automated models occasionally recommended unwarranted medication dose adjustments without physician oversight. These divergent errors demonstrate that linguistic training data deeply influences clinical safety profiles. Consequently, healthcare practitioners cannot assume uniform safety across different language settings. Providers must actively warn patients against substituting generative chatbot outputs for formal oncologic follow-up care.
To mitigate hazardous outputs, researchers formulated a structured prompt engineering strategy known as the CRAFT framework. This systematic architecture specifies context, role, audience, format, task, and tone within the conversational prompt. Subsequently, the investigators applied this refined structure to the top-performing Gemini 3.1 Pro engine. In the within-sample comparative analysis, structured prompting produced statistically significant improvements in overall clinical utility. Specifically, English response scores increased from 3.99 to 4.68 out of 5. In addition, Traditional Chinese scores advanced markedly from 4.21 to 4.84 out of 5. Most importantly, the CRAFT framework completely eliminated the unsafe advice patterns identified during baseline testing. The prompt constrained the artificial intelligence from delivering autonomous diagnostic pronouncements or unauthorized medication alterations. Instead, the optimized model effectively adopted an empathetic, evidence-based educational persona. Furthermore, this structural constraint preserved high understandability scores on the Patient Education Materials Assessment Tool. Therefore, prompt engineering provides a reproducible mechanism to control generative artificial intelligence outputs. Clinicians can utilize these standardized prompting principles when developing institutional patient portals or digital communication pathways. Structured constraints successfully bridge the gap between algorithmic capability and medical safety.
These findings deliver urgent insights for oncologists, general practitioners, and tertiary cancer care centers worldwide. Today, an increasing proportion of cancer survivors seek autonomous post-discharge guidance outside formal clinical encounters. Consequently, doctors encounter patients who bring confusing or contradictory digital recommendations into routine consultations. Clinicians must proactively explore whether patients consult online artificial intelligence tools for symptom management. Moreover, physicians must highlight that free-text conversational models lack verified clinical reasoning and accountability. In bilingual or multilingual nations such as India, these safety considerations become especially acute. Patients asking questions in regional tongues may encounter unverified traditional therapies or incorrect drug advice. Therefore, multidisciplinary cancer teams should create curated, institution-approved patient information leaflets and digital FAQ repositories. Furthermore, healthcare organizations exploring conversational artificial intelligence must mandate strict prompting architectures and multidisciplinary oversight. Regulatory bodies must also demand thorough bilingual validation before endorsing healthcare chatbot integrations. Ultimately, conversational software should assist rather than replace direct patient-provider relationships in cancer survivorship care. Clinicians remain the essential anchor ensuring safe, individualized, and evidence-based survivorship journeys.
Large language models rely heavily on their underlying linguistic training corpora. Consequently, English datasets often feature commercialized online forums and general administrative directories, leading to service routing errors and overconfident diagnostic declarations. In contrast, Chinese datasets frequently incorporate community discussions promoting traditional remedies and unverified adjunctive herbs. Therefore, chatbots mirror cultural biases and idiosyncratic management norms embedded in their regional source data, producing distinct clinical hazards across languages.
The CRAFT framework provides a standardized prompt engineering structure comprising six core elements: context, role, audience, format, task, and tone. Specifically, it instructs the artificial intelligence model to assume a restricted educational role while identifying the patient as the target recipient. Furthermore, it enforces clear formatting rules and defines exact clinical tasks while barring autonomous prescribing. Consequently, this structured approach eliminates dangerous medical hallucinations while preserving compassionate, understandable communication.
Clinicians should proactively ask cancer survivors about their digital health inquiries during every routine follow-up appointment. Furthermore, doctors must explain that public chatbots lack verified clinical liability, real-time diagnostic insight, and individualized patient context. While chatbots can assist with generic lifestyle definitions, patients must verify all symptom changes, medication questions, and supplement choices directly with their oncology team. Consequently, open clinician-patient communication shields survivors from hazardous digital medical advice.
Disclaimer: This content is for informational and educational purposes only... Refer to the latest local and national guidelines for clinical practice.
References
An H et al. Large Language Model Chatbot Responses to Cancer Survivorship Questions in Hong Kong: Bilingual Evaluation and Prompt Optimization Study. J Med Internet Res. 2026 Oct 08. doi: 10.2196/105778. PMID: 42849027.
Draelos RL, Afreen S, Blasko B, et al. Large language models provide unsafe answers to patient-posed medical questions. NPJ Digit Med. 2026;9(1):241. doi: 10.1038/s41746-026-02428-5.
Keçeci T, Karagöz B. Can large language models follow guidelines? A comparative study of ChatGPT-4o and DeepSeek. BMC Med Inform Decis Mak. 2025;25(1):350. doi: 10.1186/s12911-025-03202-5.

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A landmark study evaluates how commercial LLMs handle cancer survivorship questions across languages, exposing language-dependent safety hazards and demonstrating how structured CRAFT prompt engineering dramatically improves clinical utility.
Today

A comprehensive analysis of serum bile acid variability establishes robust reference ranges and reveals how sex, age, BMI, and smoking alter circulating profiles.
Today

A mixed-methods study evaluates a digital perinatal navigator app for high-risk pregnancies, demonstrating high usability, reduced distress, and improved awareness of vital healthcare services.
Today

A user-centered study demonstrates that AI-assisted carotid ultrasound enables nonexpert primary care staff to detect subclinical atherosclerosis. With real-time guidance and workflow optimization, this technology supports early cardiovascular risk communication and task-shifting in routine clinical practice.
Today

A hyaluronic acid-modified metal-polyphenol nanocomposite successfully eliminates reactive oxygen species in chondrocytes, halts cartilage breakdown, and promotes tissue anabolism, presenting a novel disease-modifying strategy for early osteoarthritis.
Today

A pilot randomized trial demonstrates that digital wellness applications and medically tailored meals significantly attenuate rapid weight regain following GLP-1 receptor agonist discontinuation, providing valuable transitional support for long-term obesity management.
Today