Measurement

What “scientific depth” should mean, and what it should never measure

RocketMSL5 min read
Illustration: scientific knowledge measured against a ruler

Once an organisation accepts that counting meetings tells it nothing, the obvious next step is to score the meetings instead. This is where most attempts at measuring scientific engagement quality go wrong, and they go wrong early — usually in the first design meeting, when someone proposes a measure that sounds sensible and turns out to reward the wrong behaviour.

The useful discipline is to work backwards. Before deciding what a score should count, decide what it must never count. A measure is defined at least as much by its exclusions as by its inputs, and in this domain the exclusions are where the credibility sits.

Three measures that look reasonable and are not

Duration. Longer meetings must be better meetings. In practice, length correlates with the clinician's diary, the setting, and how far the MSL travelled — not with scientific substance. Score it, and you have created an incentive to stay in the room after the useful part has finished.

Receptiveness. How positive was the clinician, how warm was the exchange, how well did it go. This is the most seductive of the three because it feels like it captures something real. It does, but not the thing anyone wants. A sceptical clinician who challenges the trial design produces a harder conversation and a worse receptiveness reading. She has also produced the more valuable interaction. Any score that rewards agreement will, over a few quarters, quietly shift field effort toward clinicians who already agree.

Self-assessment. Ask the MSL to rate the quality of their own call. This is the most common approach in practice and the least defensible, because it is unfalsifiable. It also measures reporting style rather than performance: two MSLs who ran identical meetings will score them differently according to temperament. Nobody is being dishonest. The instrument is simply not measuring what it claims to.

What these three have in common is that each measures a property of the encounter — its length, its warmth, its perceived success — rather than a property of the science that moved inside it.

The first half: was the planned science actually discussed

Start with something narrower and harder to argue with. Before an interaction, there is a set of scientific points worth raising with this clinician: the evidence relevant to her patient population, the questions she asked last time, the gaps in what she has previously engaged with. After the interaction, either those points were discussed or they were not.

Coverage of planned scientific points has two properties that the three measures above lack.

It is derived rather than imposed. The points come from the organisation's own approved scientific content, mapped to the product and the topic. Nobody sets a coverage target by judgement; the target is whatever the evidence base actually supports for this clinician. That removes the argument about whether the bar is fair.

It is falsifiable. A point was either substantively discussed or it was mentioned in passing or it was skipped. Those are different states, and distinguishing between them is a matter of evidence rather than opinion.

Coverage on its own is still incomplete, though. A conversation where the MSL worked through every planned point and the clinician said almost nothing is not a deep scientific exchange. It is a presentation.

The second half: what came back

The other half of depth is yield — what the clinician contributed that the organisation did not already know.

An unmet need described from clinical practice. An evidence gap that no publication currently fills. A treatment barrier that explains why guideline-concordant care is not happening in her institution. A question nobody had anticipated. These are the outputs that make a scientific conversation worth having at all, and they are the part that has historically evaporated into a free-text box.

Yield has to be scored for substance rather than counted. Five vague observations are not worth more than one specific, actionable evidence gap. A measure that counts insights produces more insights, most of them worthless — the same failure mode as counting meetings, one level up.

Coverage and yield together describe the two directions a scientific conversation runs in. What the organisation brought, and what the clinician gave back. Measuring scientific discussion depth as those two components, and nothing else, keeps the score pointed at the science.

What a depth score must exclude

Four things are worth capturing and worth keeping out of the score.

Sentiment. Capture it — it is genuinely useful in the clinician's profile, and a shift in sentiment over several interactions means something. Score it, and you punish the difficult conversation. These are not in tension; the same signal can be valuable as context and destructive as a metric.

The clinician's existing expertise. A world expert in the therapeutic area cannot be moved far in a single conversation, because she already knows most of what the organisation can tell her. Fold her knowledge level into a depth score and every interaction with the most valuable clinicians in the field will score badly. Knowledge level belongs in the clinician's profile, tracked over time, not inside a per-conversation measure.

Questionnaire movement. Whether the clinician says the interaction was useful is worth asking and worth recording. It is also a measure of politeness as much as of substance.

MSL competency. This is the exclusion that determines whether the field team ever trusts the measure. A depth score describes what happened in a conversation. It is not a performance rating, and the moment it is used as one, it stops being accurate — because the people producing the inputs now have a reason to shape them. The same argument, put to the field team rather than about it, is here.

The test any measure has to pass

A depth measure has to be able to say that a conversation was shallow without implying that the MSL performed badly.

Sometimes a conversation is shallow because the clinician had fifteen minutes between clinics. Sometimes because she wanted to discuss one narrow question and nothing else. Sometimes because the material available did not address what she actually needed — which is itself the most valuable thing the organisation could learn that week.

If your measure cannot distinguish those situations from poor field performance, it will be resisted, worked around, and eventually ignored, regardless of how sound the underlying model is. The design constraint is not statistical. It is that the measure has to be one the field team would be willing to defend.

RocketMSL scores coverage and yield, and deliberately excludes the rest. See how the platform measures scientific depth →