What did you ask the model to pay attention to?
“Is this normal?” and “compare the trajectory and tell me what is missing” are different task specifications even when the laboratory data are identical.
Wording shapes the task.
EXIABIO
A general-purpose chatbot does not see only your laboratory values. It also sees how you ask, what context you include and - in an ongoing chat - what came before.
The prompts below are illustrative. They are not benchmark outputs from a named AI product, and this Brief does not claim that every model will respond in the same way.
Prompt wording can foreground reassurance, trajectory, missing context or causal restraint. A knowledgeable user may ask for those safeguards explicitly. Most consumers should not have to know the perfect scientific prompt before a health-AI system behaves responsibly. [1-5]
Everything is still normal. Is there anything to worry about?
The wording foregrounds “normal” and reassurance. It does not explicitly request longitudinal comparison, missing-context checks or causal restraint.
Compare these three tests over time. Flag important directional changes, keep missing context explicit, and do not assume a cause.
This specifies a more disciplined task - but better prompting still does not turn a general-purpose chatbot into a validated interpretation system.
In patient-facing use, questions are often incomplete, colloquial or framed around a prior belief. Research also shows that multi-turn conversation can influence model behavior in ways that are not obvious from the current question alone. [1-4]
“Is this normal?” and “compare the trajectory and tell me what is missing” are different task specifications even when the laboratory data are identical.
Wording shapes the task.A user's belief, an earlier AI hypothesis or an irrelevant prior turn can become additional context for the next response unless its status stays explicit.
History can become hidden input.Relevant history can make an answer more useful. But multi-turn medical studies have also observed models shifting toward a user's incorrect position, degrading across sequential alternatives, or propagating earlier model errors into later turns. [2-4]
That does not make every history-dependent answer a “hallucination.” A broader problem is provenance collapse: measured facts, user beliefs and generated hypotheses can start to look like the same kind of truth.
Laboratory facts, user-reported context, model hypotheses and external evidence can appear in one conversation. Their provenance should stay visible.
Values, units, dates and other recorded laboratory information. These still need provenance and comparability checks.
Symptoms, medication use, lifestyle changes or beliefs supplied by the person can matter while remaining user-reported information.
An earlier AI sentence such as “this may be related to X” should not silently become established history in the next turn.
Scientific evidence should be distinguishable from conversational claims and should fit the specific conclusion being made.
A consumer should not need to know the perfect prompt
before a health-AI system behaves responsibly.
Prompting can improve task specification. Governance should decide what the system must check even when the user does not know to ask.
Go deeper: Can AI Reliably Interpret Blood Test Results? →These sources support the prompt/context and multi-turn reliability propositions used here. They do not show that every chatbot, model version or conversation will fail in the same way.