Skip to content
Arts in Health Institute

Music, art, dance and drama in health care: what the trials actually show, what is still only promising, and how people get referred.

Music, art, dance and drama, and the health claims made for them.

Reading Arts in Health Research: Effect Sizes, Certainty and Where a Figure Came From

Published · Last reviewed

Before asking what a number in arts in health says, ask what kind of study produced it, because in this field the two most commonly quoted figures on the same page are often a randomised trial and an uncontrolled evaluation, presented at identical volume. Everything else on this page follows from that one habit. Every efficacy claim on this site carries a tier label (Tier 1 supported, Tier 2 promising but limited, Tier 3 practice and anecdote), and this is the page that explains what the labels are doing.

I started reading this literature because of an afternoon in my father’s care home, not because of an interest in methodology. What changed my mind about method was a funding meeting a few years later, in which a very good and entirely sincere project manager told the room that the National Institute for Health and Care Excellence recommends music therapy in dementia. I went home and searched the guideline. Music, art, dance, drama and creative activity return zero occurrences in the whole of NG971. Nobody in that room was lying. The claim had simply been passed along enough times that checking it had stopped seeming necessary. This page is the set of checks I now run before I repeat anything.

If you want the ladder itself rather than the reasoning behind it, the editorial policy sets out what a claim has to show to earn each tier.

What counts as controlled evidence, and what does not

Controlled evidence compares the people who got the activity against people who did not, or who got something else. Without that comparison, a study cannot tell you whether the activity did anything.

This matters more in arts in health than in most fields, because the sector produces a very large number of before and after evaluations and almost no trials. The 2024 systematic review of arts on prescription screened 7,805 records, included 25, and found that 10 were qualitative, 7 were mixed methods and 8 were quantitative studies using uncontrolled pre-post designs. Its own words: no randomised controlled trials were identified in the search2. That review did report a statistically significant improvement in wellbeing when it pooled the quantitative work, and that pooled figure rests entirely on studies with nothing to compare against, which is why it stays Tier 3 here.

An uncontrolled before and after design cannot separate the activity from four other things: time passing, regression to the mean (people usually join when they are at their worst, and the worst point tends to be followed by a better one whatever happens next), who chose to take part, and who dropped out before the second measurement. None of that makes the evaluations worthless. They describe what a service did, who came, and what people said about it, and this site quotes plenty of them. It just never calls them trials.

Volume of studies is not strength of evidence

A review can pool a great many trials and still conclude that it cannot trust the answer, and the clearest example in this whole field is worth memorising.

The 2021 Cochrane review of music interventions in people with cancer included 81 trials and 5,576 participants. It reports anxiety at 7.73 STAI-S units lower and pain at SMD -0.67, and it rates both findings very low certainty3. That is Tier 2 on this site: controlled evidence exists, and it is not strong enough to lean on.

Eighty one trials. Five and a half thousand people. Very low certainty. If you understand why those three facts sit together comfortably, you have understood most of what this page is trying to teach. Certainty ratings ask how much confidence the pooled estimate deserves, given how the individual trials were run: whether they were small, whether they were at risk of bias, whether they agreed with each other, whether the outcome was measured in a way that could be nudged. Eighty one small, unblinded, inconsistent trials produce a precise looking average built out of weak material. The study count is the least informative number in an abstract, and it is the one that ends up in the press release.

What an effect size tells you that a p value does not

A p value tells you how surprising a result would be if there were no real difference. An effect size tells you how big the difference is. Only the second one is any use for deciding whether something is worth your Tuesday afternoon.

Take the 2025 Cochrane review of music based therapeutic interventions in dementia, which covers 30 studies and 1,720 participants randomised. For depressive symptoms it reports SMD -0.23 (95% CI -0.42 to -0.04) at moderate certainty4. That is Tier 1: a real effect, from evidence the reviewers trusted reasonably well, and a small one. Roughly a fifth of a standard deviation. It is not nothing and it is not transformation, and a page that reported only “significant improvement in depressive symptoms” would have told you the true thing and left you with the wrong impression.

Effect sizes also travel badly. An SMD is a difference between group averages. Inside the trials producing that -0.23, some people improved considerably and some did not improve at all, and nothing in the figure tells you which one a particular person will be. That caveat applies to every Tier 1 claim on this site, including the strongest one.

A confidence interval that sits across zero

An interval spanning zero means the data are compatible with a benefit, with a harm, and with nothing. Whether that counts as an open question or a closed one depends on how narrow the interval is and how much the reviewers trusted it.

The same dementia review reports agitation and aggression at SMD -0.05, 95% CI -0.27 to 0.17, at moderate certainty4. Agitation is the outcome most often claimed for music in dementia, which makes this the most important number on the site. The interval is narrow, it is centred essentially on zero, and the certainty rating is moderate rather than low. Read together, that is Tier 1, null result: moderate certainty evidence of no meaningful effect on agitation. It is a positive finding of absence, not a hole in the literature, and reporting it as “more research is needed” would misdescribe it.

The same review found no evidence of any effect persisting four weeks after the sessions end, which is Tier 3, and it contains no separate quality of life estimate at all, so any quality of life figure attributed to it has come from somewhere else. The detail sits on what the dementia evidence actually shows.

Whole sample first, subgroup second

When a trial’s overall result is null and a subgroup result is positive, the honest order is overall result first. Reversing that order can be perfectly accurate sentence by sentence and still leave a reader with a false picture.

The 2018 three arm trial of singing for postnatal depression is the cleanest illustration I know. Across the whole sample there was no significant effect (p=0.16). The significant finding is in a moderate to severe subgroup at week 6, and it was no longer significant by week 10. The trial had an active comparator, creative play, and singing did not beat it5. The fair summary is faster recovery among the more severely affected women, not a better endpoint, and it is Tier 2 precisely because it rests on a subgroup at an early timepoint.

Watch for four subgroup tells: a result described as appearing “particularly in” some group; an outcome reported at one timepoint when several were measured; a comparison against usual care when an active comparator was also in the trial; and a headline about a trial whose primary outcome is never actually named.

Certainty ratings, and the reviews that predate them

Not every review carries a certainty rating, and the absence of one is not a low rating. It is usually a date.

The Cochrane review of music and preoperative anxiety includes 26 trials and 2,051 participants and reports anxiety 5.72 STAI-S units lower (95% CI -7.27 to -4.17)6. It is Tier 1 here. But it was published in 2013, before certainty assessment became routine in these reviews, so it carries no certainty rating and none can be invented for it. Describing it as high certainty or moderate certainty would be attributing to the review a judgement the review never made. The honest move is to give the trial count, the sample size and the interval, and to say why there is no rating. More on that unusually clean result at music before surgery.

Units, and the number that must not be converted

Quote a figure in the units its source used. Conversion is where numbers quietly get better.

The best evidenced finding in this whole field is rhythmic auditory stimulation after stroke: gait velocity 11.34 m/min faster (95% CI 8.40 to 14.28), from 9 trials and 268 participants, at moderate quality7. Tier 1, and unambiguously so. The review reports metres per minute. Re-expressing it in centimetres per second produces a different looking number and no new information, and once one figure on a page has been converted for effect the reader has no way of knowing which others have. Note also how narrow the claim is: gait velocity, and nothing about mood, communication or quality of life. The full picture is on rhythmic auditory stimulation.

How to read a nested count

When a document reports several counts of studies, check whether they are additions or containers before adding them up.

The WHO Europe scoping review is the most misquoted example. It covers over 900 publications, which comprise 200 plus reviews and 700 plus individual studies; those reviews between them cover over 3,000 studies8. So the 3,000 sit inside the 900, not alongside them. Written as “3,000 studies and 900 publications”, the same document appears several times larger than it is. And a scoping review maps a literature rather than pooling it, so it carries no tier and supports no claim about effect. It is evidence that a field has been studied. That is a real and useful thing to know, and it is not the same as evidence that anything in it works. Policy reports and what they are for goes further into that distinction.

Trace the figure to the document it came from

If a striking number is quoted without a study behind it, the number is usually doing a different job than the one it appears to be doing. Go to the original.

The savings figures for arts on prescription are the case I would send anybody to first. They trace to a cost benefit summary of the Artlift scheme in Gloucestershire, written by a GP in December 2011, which describes itself in its own conclusion as “a simple observational study” and states plainly that it “does not imply causality”9. Read the document and three things emerge that the quotations lose:

  • The 37% fall in GP consultations is the months 7 to 12 figure, from 11.3 consultations a year to 7.1. The report’s own full year figure, printed in the same paragraph, is 24%.
  • The 27% is a reduction in overall NHS spend (£157,473 before, £115,050 after), not in admissions. Admissions appear separately as 54 before and 33 after, with no percentage and no significance test attached, which is probably how the “fewer admissions” version got started.
  • The per patient saving is assembled from a £471 figure, which is £42,423 divided by the 90 patients whose spend was analysed, and a cost of £360, which is £180,000 divided by the 500 patients referred. The numerator and the denominator come from different populations.

The sample was 90 patients out of roughly 500 referred over three years, with no control group. That makes the whole cluster Tier 3, and the 2024 finding of no randomised trials anywhere in this literature2 is why it stays there rather than being a contested Tier 2. The full trail is set out on does arts on prescription save money, and the broader context on arts on prescription.

Two shorter examples of the same habit. Ulrich’s 1984 window study, still cited in hospital design documents four decades on, is 46 patients in 23 matched pairs, one Pennsylvania hospital, one operation, one season, retrospective records from 1972 to 198110, which is Tier 3 however often it is repeated: see does hospital art speed recovery. And a number can turn out not to be a statistic at all. The credential count widely quoted for art therapists in the United States comes from a banner attached to a portal transition notice. The board’s own 2021 annual report gives a headline total of 6,235 whose Active rows sum to 4,95511; its 2020 report gives 8,226 where the rows sum to 8,053 and Active to 7,15012; and the count of active registered art therapists drops from 3,165 to 1,861 between the two with no explanation offered. The credentials also stack, so none of those tables counts people. This site therefore publishes no US art therapy credential figure at all, which is covered on art therapy.

What a Tier 3 label means, and what it does not

Tier 3 means either that nobody has run an adequately powered controlled study, or that the controlled evidence that exists is null. Those are different situations and the article always says which one it means.

This matters because the error runs in both directions. Overclaiming is the obvious failure, and this field does plenty of it. The opposite failure, treating “not properly studied” as “shown to be useless”, is just as misleading and does more damage to the people it lands on. Most of what happens in care home lounges and community halls sits at Tier 3, including work I do myself and would defend. What I cannot do is promise you it will help, and what nobody can honestly tell you is that it will not. Dance and depression is a good test case: the Cochrane review there covers 3 studies and 147 participants and reports SMD -0.67 (95% CI -1.40 to 0.05) at very low quality, with its authors declining to draw firm conclusions, while the number that actually circulates from it is quoted with the wrong label attached and is not an effect size at all.

A short checklist

Before repeating a figure from anywhere, including from here:

  1. What design? Randomised, controlled but not randomised, or before and after with no comparison.
  2. How many people, and how many studies? Both, and in that order of importance.
  3. What is the effect size, and what is the interval? Not just whether it was significant.
  4. Is there a certainty rating? If not, is that because the review predates them, or because nobody has said?
  5. Whole sample or subgroup, and at which timepoint?
  6. Is this the current version of the review?
  7. Which outcome, exactly? A result about gait velocity is not a result about mood.
  8. Can I open the source? If the trail ends at a report quoting a report, the figure has not been checked, it has been forwarded.

That is the whole method. It takes about ten minutes per claim, which is longer than sharing a clip and shorter than a course of therapy that was never going to help.

Frequently asked questions

What is the difference between a trial and an evaluation?

A trial allocates people to the activity or to something else, so that the two groups can be compared. An evaluation usually measures one group before an activity and again afterwards, with nothing to compare against. That second design cannot separate the effect of the activity from time passing, from regression to the mean, or from the fact that people who sign up for a ten week art course are different from people who do not. Both are useful documents. Only the first can support a claim that the activity caused the change, and on this site an evaluation is never called a trial, however the report quoting it describes itself.

What does an effect size like SMD -0.23 actually mean?

SMD stands for standardised mean difference, and it is used when the pooled trials measured the same idea on different scales. It expresses the gap between groups in standard deviations rather than in points on any one questionnaire. An SMD of -0.23 is roughly a fifth of a standard deviation, which is conventionally described as a small effect. It is a difference in group averages, not a prediction about a person: within any trial reporting it, some participants improved a great deal and some got worse.

Why would a review with 81 trials still be rated very low certainty?

Because certainty is about how much the pooled result can be trusted, not about how much of it there is. The 2021 Cochrane review of music interventions in cancer care included 81 trials and 5,576 participants and rated its anxiety and pain findings very low certainty, chiefly because the individual trials were small, at high risk of bias, and inconsistent with one another. Pooling a large number of weak studies produces a precise looking number built on weak inputs. Study count is the least informative figure in any abstract.

Does 'no evidence' mean the activity does not work?

No, and the two get confused constantly in both directions. 'No adequately powered controlled evidence' is a statement about the research literature. 'It does not work' is a statement about the world. A great deal of arts in health has never been properly studied for reasons unconnected to whether it helps: the interventions are hard to blind, funding is short term, and the sector's money goes into delivery rather than trials. Where the controlled evidence exists and is genuinely null, this site says null in those words, because that is a different and stronger statement than an empty file.

What does it mean when a confidence interval crosses zero?

It means the data are compatible with an effect in either direction, including none at all. What matters next is how wide the interval is and how certain the reviewers were. The 2025 dementia review reports agitation and aggression at SMD -0.05 with an interval of -0.27 to 0.17, at moderate certainty. That is a narrow interval sitting around zero, from evidence the reviewers trusted reasonably well, so it reads as a finding that music based interventions do not meaningfully change agitation, rather than as an unanswered question.

How do I check I am reading the current version of a Cochrane review?

Look at the pub number at the end of the identifier and the publication date on the record itself, not the date on the article quoting it. The dementia review is CD003477.pub5, published 7 March 2025, covering 30 studies and 1,720 participants randomised. Its predecessor covered 22 studies and 1,097 participants, and reached different conclusions on agitation. Any source quoting the smaller numbers today is quoting the superseded version, and that includes sources published well after the update.

Why does this site refuse to convert figures into other units?

Because converting is where flattery creeps in. A gait velocity result reported as 11.34 metres per minute can be re-expressed in centimetres per second, and the resulting number looks quite different without anything having been learned. The same goes for turning a small standardised effect into a percentage, or a per patient average into an annual total. Quoting a figure in the units its source used is the cheapest available protection against accidentally improving it.

References

  1. Dementia: assessment, management and support for people living with dementia and their carers (NG97), National Institute for Health and Care Excellence.
  2. The impact of arts on prescription on individual health and wellbeing: a systematic review with meta-analysis, Jensen A, Holt N, Honda S, Bungay H, Frontiers in Public Health, 9 July 2024.
  3. Music interventions for improving psychological and physical outcomes in people with cancer, Cochrane Database of Systematic Reviews, CD006911.pub4, 2021.
  4. Music-based therapeutic interventions for people with dementia, Cochrane Database of Systematic Reviews, CD003477.pub5, 2025.
  5. Effect of singing interventions on symptoms of postnatal depression: three-arm randomised controlled trial, Fancourt D, Perkins R, British Journal of Psychiatry, 2018 (PMID 29436333).
  6. Music interventions for preoperative anxiety, Cochrane Database of Systematic Reviews, CD006908.pub2, 2013.
  7. Music interventions for acquired brain injury, Cochrane Database of Systematic Reviews, CD006787.pub3, 2017.
  8. What is the evidence on the role of the arts in improving health and well-being? A scoping review, WHO Regional Office for Europe, Health Evidence Network synthesis report 67, 2019.
  9. Cost-benefit evaluation of Artlift 2009-2012: summary, Dr Simon Opher, 9 December 2011.
  10. View Through a Window May Influence Recovery from Surgery, Ulrich RS, Science, Vol. 224, No. 4647, 27 April 1984.
  11. ATCB Annual Report 2021, Art Therapy Credentials Board.
  12. ATCB Annual Report 2020, Art Therapy Credentials Board.

Written by Miriam Halstead. Reviewed by Dr Anna Bergström, PhD, MSc Epidemiology.

Our guides are written from personal experience and reviewed by a registered arts therapist for accuracy. Read our editorial policy.

Related articles