Empathic AI for Cancer Patients

This project introduces a framework for iteratively tuning and evaluating empathy in large language model (LLM) responses, demonstrated through a conversational agent for cancer patients. The approach combines prompt engineering for empathy tuning with automated evaluation and user feedback for measurement. Responses are grounded in domain knowledge retrieval (RAG) to ensure they are both emotionally supportive and factually accurate.

Text
Illustration of an empathic interaction between AI and a patient generated by DALL-E

Initial situation

In oncology, patients face complex treatments and major emotional challenges. A cancer diagnosis often brings fear, uncertainty, and isolation. Empathy is vital for patient-centred care, improving satisfaction, trust, and treatment adherence. Yet healthcare professionals frequently have limited time to offer sustained emotional support. This can create communication gaps, leaving patients feeling unheard or reduced to a diagnosis. With the rise of AI, there is potential to deliver scalable, consistent, and context-aware empathic support that complements human care, if they are carefully designed and evaluated to foster genuine emotional connection.

Problem statement / Project goal

This work addresses the challenge of how to precisely tune and measure empathy in LLM-based systems for emotionally sensitive healthcare contexts.
The project goals were to:

  • Develop an iterative empathy tuning and measurement process combining prompt engineering, EPITOME-based evaluation, and real-time user feedback.
  • Ground responses in verified oncology information via retrieval-augmented generation.
  • Examine how prompt design and communication strategies influence both system-expressed and user-perceived empathy.
  • Validate the framework with non-patients, a cancer patient, and members of a cancer support organisation.

Conversational Agent

FHNW

Chat interface providing fact-based answers and emotional support to cancer patients, grounded in verified oncology knowledge for both empathy and accuracy.

Framework

FHNW

Visual overview of the iterative loop for tuning and measuring empathy in AI responses, integrating prompt engineering, automated EPITOME scoring, and user feedback to refine both system-expressed and user-perceived empathy.

Empathy Dashboard

FHNW

Monitoring interface displaying EPITOME scores and user feedback, enabling iterative prompt improvements while maintaining a balance between empathy and factual accuracy.

Solution developed and its benefits

We developed and implemented a framework for tuning and measuring empathy in LLM-generated responses, demonstrated through a conversational agent for cancer patients. While designed for healthcare, the framework is domain-agnostic and can be applied to any setting where trust, accuracy, and emotional support matter.

The framework works as a continuous improvement loop with two tightly connected parts:

Empathy Tuning

The tuning module shapes the chatbot’s behaviour through structured prompt design. Each prompt defines the tone, style, and level of empathy while ensuring factual alignment with verified oncology knowledge retrieved via RAG. Prompts are versioned so their performance can be compared across iterations.

Empathy Measurement

Every response is evaluated automatically using the EPITOME framework, which scores emotional reactions, interpretations, and explorations, and in parallel through user Likert ratings collected directly in the chat interface. These two perspectives, system-expressed empathy and user-perceived empathy, are reviewed together in an administrative dashboard, guiding targeted prompt refinements.

By combining these two processes in a single loop, the framework enables measurable, data-driven refinement of empathy in AI communication.

Benefits

  • Operationalises empathy as a measurable, tunable quality in AI responses.
  • Captures both system-expressed empathy and user-perceived empathy.
  • Enables comparison of results and iterative refinement of empathic responses.
  • Maintains factual reliability via RAG, preventing misinformation in sensitive domains.
  • Transferable beyond oncology to any setting requiring emotionally attuned AI communication.

Key Findings

Applying the framework in controlled tests and real-world trials revealed important insights into how empathy can be shaped and experienced in AI communication. Structured prompts with clear, step-by-step guidance consistently produced higher EPITOME scores than generic empathy instructions, confirming that LLMs require explicit procedural direction to express empathy reliably. However, high automated scores did not always align with how users experienced the responses. Many valued moderate, context-sensitive empathy over highly emotional phrasing, and preferred reassurance delivered before factual information.

In interactions with cancer patients, especially those with lived experience, scripted or overly generic phrases were sometimes perceived as insincere or burdensome. Direct, needs-oriented questions often felt more supportive. Among older participants, adoption was limited by low digital literacy, privacy concerns, and a preference for human contact over AI interaction.

These findings show that while empathy can be tuned and measured systematically, its real-world effectiveness depends on tone, timing, authenticity, and the user’s trust in and willingness to engage with the system, factors that must be considered alongside technical optimisation.

Key terms

  • Empathy TuningIterative process of adjusting LLM prompts and strategies to influence the type and quality of empathy expressed in responses.
  • Empathy MeasurementUse of automated evaluation and user feedback to assess how well a system expresses empathy and how it is perceived by users.
  • EPITOMEEmpathy Process Task-Oriented Model Evaluation; A structured framework for measuring empathy in generated responses in terms of Emotional Reactions, Interpretations, and Explorations.
  • Retrieval-Augmented Generation (RAG)Integrates a retrieval system with an LLM to ground responses in verified domain-specific knowledge.
  • Conversational AgentAn AI system designed to engage in dialogue with users via natural language, here used as the case study for applying the empathy tuning and measurement framework.

Customer

IMPortant Logo

Horizon Europe - Project IMPortant
Visit the Project’s official website for details

Team

Daniel Barber
Tamira Leber

Advisors
Prof. Dr. Samuel Fricker
Marjam Yeganeh