Insights

"Evidence-based" is a big claim. Here are four questions to test it

A plain-language guide for anyone who funds, runs or joins a health or wellbeing program

Illustration: A brass magnifying glass resting on an open book of blank cream pages, with a small burgundy ribbon marking one page

Imagine you are reading a grant page. (This is an imagined page, but you may have seen ones like it.) It says: "Our program is evidence-based." The words are calm and confident. You want to believe them.

What would help you decide?

"Evidence-based" is a good phrase. It points at something we hope for: that care and money go to what is actually likely to help. The trouble is that the phrase can mean very different things. It can mean "a large, careful study found this works." It can also mean "we read some research and it inspired us."

Both are honest starting points. They are not the same, and a reader deserves to know which one is on the page.

A ladder, not a stamp

The UK innovation foundation Nesta describes a framework called the Standards of Evidence in a 2013 paper. It does not treat evidence as a yes-or-no stamp. It treats it as five steps, from a clear and logical account of why a program could help, up to evidence that it keeps working when someone else runs it somewhere else. Nesta's stated goal is to help us "know how confident we can be in the evidence."

It also does something kind. It does not demand the top step from every new idea. Its first step is described as appropriate for early-stage ideas, and the framework says it does not demand particular methods.

We find that hopeful. A young program does not have to pretend. It only has to be honest about where it stands.

The four questions below are a simplified guide inspired by that ladder; they are not Nesta's five-level assessment.

Question 1: What changed, and was it measured?

There is a difference between telling a good story about a program and collecting data on what happened to the people in it. Nesta's second step is gathering data that "shows some change amongst those receiving or using" the program. Nesta is clear about the limit of this step: it shows change, but it does not show that the program caused it.

So ask: Did they measure something about the participants before and after? What exactly? A reasonable answer might be a short, named survey. A weak answer is only a testimonial.

Question 2: Compared with whom?

People change for many reasons. They get older, the season turns, they find a new job, they feel better after a hard time. So a before-and-after change may have happened without the program at all.

Nesta's third step is where this gets tested. It asks for evidence of causing the impact by "showing less impact amongst those who don't receive" the program. In its table, Nesta describes methods using a control group, and says random selection of participants strengthens the evidence.

So ask: Is there a comparison group of similar people who did not get the program? If there is not, the claim may still be reasonable, but it is not at this step yet.

Question 3: Who checked, and was everything shared?

A program that checks its own work has an obvious reason to hope for good news. That is human, and it is why Nesta's fourth step asks for an independent evaluation.

There is a second half to this question, and it is about what gets published. In 2008, researchers compared what the US Food and Drug Administration held on 74 antidepressant studies with what appeared in journals. About 31% were unpublished in the researchers' review. According to the published literature, it appeared that 94% of the trials were positive. In the FDA's own analysis, 51% were. The researchers wrote that "evidence-based medicine is valuable to the extent that the evidence base is complete and unbiased." They also noted they could not tell whether authors, sponsors or journals were mainly responsible.

That is one field, studied once, so it would be wrong to say it describes every program. But it shows why this question matters.

So ask: Did someone with no stake in the result do the evaluation? Are all the results available, including the disappointing ones?

Question 4: Would it work somewhere else, run by someone else?

A program can work wonderfully where it began, run by the people who dreamed it up, and look quite different in a new place. Nesta's fifth step asks whether a program can be run by someone else, somewhere else, while continuing to produce a positive effect.

A large project in psychology shows why one good result is a starting point. Researchers attempted to replicate 100 experimental and correlational studies published in three psychology journals. Ninety-seven percent of the originals had statistically significant results. Thirty-six percent of the repeats did, and the repeated effects averaged about half the size of the originals. That does not mean the originals were worthless. It suggests that one study, however good, is the beginning of an answer.

So ask: Has anyone else tried it, in a place like ours? Is it written down well enough that another team could run it?

What these questions are not

They are not a way to dismiss a program. Nesta itself says that even a program at its top step is "not an end point," because evidence "may only ever be partial or timebound."

They are also not a test that every community program must pass in a year. A small, new group doing something thoughtful may honestly sit on the first or second step. That is fine. The only thing that is not fine is a label that suggests the fifth step when the evidence sits at the first.

A good answer often sounds like this: "We are at step two. Here is what we measured. Here is what we do not yet know. Here is how we plan to find out."

The world we are working toward

Illustration: A ladder of five simple wooden steps rising toward a window of warm light, in a cream room with a burgundy rug

Imagine a world where saying "evidence-based" always came with a small note: which step, measured how, checked by whom. Where a grant page said "we do not know this yet" and funders thought better of the program for saying so.

Illustration: Five small brass weights of different sizes laid in a row on a cream cloth beside an old balance scale

Imagine that habit spreading to a clinic in a mountain town, a school in Guatemala City, a community center in Pune, and not only to organizations with an evaluation department.

That is the world the Global Wellbeing Institute is building toward. Our Impact & Evaluation Lab, now in development, is planned as GIHEW's research and outcomes-measurement program. We hope to hold ourselves to these same four questions.

Illustration: A village path at dusk leading toward a cluster of warm lit windows, each slightly different

If you fund, run or take part in community wellbeing work, and you would like to see this kind of plain honesty spread, we would love to hear from you. Come and build it with us.

Request our prospectusdirector@globalwellbeinginstitute.org

Sources

Global Wellbeing Institute is a new organization and its programs are in development. We do not report program results.