Field medical

It measures the conversation, not you

Zainab Zahid, Head of Product5 min read
A luminous data waveform running between two faint human figures — the conversation itself, rather than either person in it

MSL performance measurement usually means activity targets and a self-rated call quality score. Scientific Discussion Depth is a different kind of measure: it scores what happened in the discussion, and it is designed so that a difficult conversation scores well rather than badly.

If you run field medical meetings for a living, you have seen a measurement tool arrive before. It is introduced as support, it is described as a productivity gain, and within two quarters somebody is being asked why their numbers are lower than a colleague's.

So the reasonable assumption, when a platform says it measures scientific conversations, is that it is measuring you. I want to explain why that is not what this is, in terms that are structural rather than reassuring — because a promise is worth very little and a design constraint is worth quite a lot.

Key takeaways

  • Scientific Discussion Depth scores the conversation that took place, not the Medical Science Liaison who conducted it
  • The score is built from talking-point coverage and insight yield, and excludes sentiment, HCP receptiveness and any assessment of the MSL
  • A challenging healthcare professional produces a harder conversation and a better score, not a worse one
  • MSLs are not ranked against one another anywhere in the product
  • Nothing is recorded without HCP consent, captured per meeting, and declining consent does not penalise the interaction

Why does every measurement tool feel like surveillance?

Because most of them are, structurally, even when nobody intended it.

A tool that captures what happened in your meetings and reports it upward has created a performance record whether or not it was designed as one. The intent of the person who bought it does not change what the data can be used for eighteen months later, when a different manager is running the team.

That is a legitimate concern and it is not solved by a reassuring paragraph in a rollout deck. It is solved, or not, by what the measure is made of. So it is worth being specific about that, and about why the function needs a measure at all.

What is actually being scored?

Two things, and deliberately only two.

Talking-point coverage. Before a meeting, there is a set of scientific points worth raising with that particular HCP — drawn from your own approved content, the questions they asked last time, and the topics they have not previously engaged with. Afterwards, the score reflects how many of those were substantively discussed. Not mentioned in passing. Discussed.

Insight yield. What the HCP contributed that the organisation did not already know. An unmet need from their practice, an evidence gap, a barrier to guideline-concordant care, a question nobody had anticipated.

That is the whole score. Everything else that gets captured during an interaction — sentiment, receptiveness, how the HCP responded to you — is recorded for the profile and excluded from the number. There is a longer argument for what a depth score should exclude and why, but the short version is that including any of it would make the measure worse.

What happens to a hard conversation?

This is the question that matters most, so here is the direct answer: it scores well.

Take the meeting nobody wants. A sceptical academic challenges the trial design, questions whether the primary endpoint applies to her patient population, raises a comparator you would rather not discuss, and ends the conversation unconvinced.

On sentiment, that is a poor meeting. On any receptiveness measure, it is a poor meeting. On a self-rated call quality score, most people would mark it down out of honesty.

On Scientific Discussion Depth it is a strong meeting, because a large amount of planned science was substantively engaged with and the HCP produced three specific evidence gaps the organisation did not previously know existed. Those evidence gaps are worth more to the medical plan than a dozen agreeable conversations.

This is not generosity in the scoring. It is the consequence of excluding sentiment. A measure that rewards agreement will always, eventually, push a field team toward the HCPs who already agree — and everyone loses, including the HCPs who most needed the harder conversation.

What the platform will never do

Four commitments, stated plainly because vagueness here is worthless.

MSLs are not ranked against each other. There is no leaderboard. The score belongs to the conversation, and conversations are not comparable across territories, therapeutic areas or HCP populations in any way that would make a ranking meaningful.

Nothing is recorded without HCP consent. Consent is captured for each meeting, not once for a relationship. If it is declined, no audio is captured and nothing is transmitted anywhere. You complete a structured record instead, and that interaction is not marked down for lacking a recording.

No new forms. This replaces post-meeting admin rather than sitting on top of it. If you find yourself entering the same information twice, that is a defect and it should be reported as one.

You correct the AI, not the other way round. Every generated insight, summary and coverage assessment is reviewed before it enters the record. If the model has misread what an HCP said, you change it, and your version is the record.

What this changes about your week

The honest answer is: mostly the admin.

Preparation stops being an evening spent assembling context from four systems, because the context is already assembled. The post-meeting write-up stops being composition and becomes review and correction, which takes a fraction of the time and is considerably less unpleasant at seven in the evening.

What it does not change is the conversation itself. The judgement about which point to press, when to stop talking, and what a particular clinician actually needs is the part of the job that is not automatable and is not being measured as though it were.

And there is one further change worth naming. When your function can evidence what it delivers, the argument for resourcing it gets stronger — which affects headcount, evidence generation budgets, and whether Medical is in the room for decisions it currently hears about afterwards. That is worth something to the people doing the work, not only to the people reporting on it. It is also the opposite of how activity targets reshape a diary.

See how the platform works →

Zainab Zahid is Head of Product at RocketMSL.