Understanding how nutrition evidence is graded, and why the rules built for drug trials don’t always fit

Written by Steven Paul, Registered dietitian


One week eggs are back on the menu. The next week a headline says otherwise. One study links red meat to disease risk, then another says the risk is smaller than the coverage suggested. If you’ve ever asked yourself why does nutrition advice keep changing so often, you’re far from alone.

A good part of the answer sits inside the research itself. Nutrition science borrows its rules for judging evidence from drug research, and those rules were built to answer a different kind of question. That mismatch is often where the confusion starts, more than any single contradictory headline.

Nutrition advice does not really change as often as it feels like it does. What varies is the strength of evidence behind different claims, and scientists are still genuinely divided over how best to grade that strength. There is no single, universally agreed system for rating nutrition evidence, and the disagreement between the main systems is real, published and unresolved. That isn’t a sign nutrition science is broken.

Why can’t nutrition research just run drug-style trials?

In medicine, a randomised controlled trial (RCT) is often treated as the strongest evidence available. Participants are randomly assigned to different groups so that, before the study starts, the groups are as similar as possible in every way except the thing being tested. Drug trials usually add blinding too, where neither participant nor researcher knows who received the real treatment and who received a placebo, an inert substance used for comparison. This stops expectation alone shaping the result.

Diet doesn’t work like that. You can’t hide from someone that they’re eating a Mediterranean-style diet instead of their usual meals, the way you can hide whether a tablet is active or inert. Ludwig, Ebbeling and Heymsfield made this point in a 2019 JAMA perspective, arguing that dietary assignment can rarely be properly blinded, and that nutrition research needs its own design standards rather than being judged purely against a template built for drugs. This is one research group’s argued position, not a settled fact, but it captures a structural problem most researchers in the field accept exists in some form.

When do randomised trials work well in nutrition?

Randomised trials do work well in nutrition, particularly in short, tightly controlled experiments where researchers can measure exactly what someone eats.

Hall and colleagues demonstrated this in a 2019 Cell Metabolism study. Twenty weight-stable adults spent two weeks on an ultra-processed diet and two weeks on an unprocessed diet, in random order, with both matched for the calories, sugar, fat and fibre on offer, and were free to eat as much or as little as they wanted. On the ultra-processed diet, participants ate around 508 kilocalories more per day on average, and gained roughly 0.9 kilograms over the two weeks, while losing a similar amount on the unprocessed diet.

This is a genuinely well-controlled experiment, and it tells us something real about short-term appetite and intake. It’s also worth being honest about its limits; twenty people, four weeks in total, inside a research unit, is a long way from someone’s everyday life over months or years.

What happens when a big trial contradicts years of observational research?

Some of the most useful lessons in nutrition science come from moments when a large randomised trial produced a result running directly against earlier, non-randomised research.

The clearest example is beta-carotene. Observational research had linked diets rich in beta-carotene with lower cancer risk. When researchers tested this directly in the CARET trial, published by Omenn and colleagues in 1996, the result went the other way; among more than 18,000 participants at high risk of lung cancer, those given beta-carotene and retinol supplements had more lung cancers and more deaths than those given a placebo. The trial was stopped early. It remains the case researchers reach for when explaining why observational associations need testing before anyone treats them as proven.

The Women’s Health Initiative Dietary Modification Trial, published by Howard and colleagues in 2006, tells a different kind of cautionary story. Almost 49,000 postmenopausal women were randomly assigned to their usual diet or a low-fat pattern and followed for just over eight years on average. The low-fat pattern did not significantly reduce heart disease, stroke or overall cardiovascular disease, partly because adherence slipped; the gap in fat intake between groups narrowed over time as people drifted back toward old habits, a recurring problem in trials asking people to sustain a different diet for years.

A third example cuts the other way. The PREDIMED trial, originally published by Estruch and colleagues in 2013, tested a Mediterranean-style diet against a lower-fat comparison diet in people at cardiovascular risk. It was retracted and republished in 2018 after researchers found randomisation problems at a handful of sites, affecting around 1,588 of the original 7,447 participants. Once the data were reanalysed correctly, the headline finding, roughly a 30% relative reduction in major cardiovascular events, held up. Even a well-known, influential trial can have real methodological problems, and the right response is careful correction rather than blind trust or blanket dismissal.

Do observational studies and randomised trials reach the same conclusions?

A cohort study follows a large group of people over time, records what they eat and what happens to their health, and looks for patterns, without assigning anyone to a particular diet. A natural question is whether this kind of evidence broadly agrees with randomised trials on the same topics.

Schwingshackl and colleagues tried to answer this directly in a 2021 BMJ study, matching 97 pairs of nutrition questions where both trial-based and cohort-based evidence existed. On average, the two types of evidence lined up fairly closely, though the range of results across individual comparisons was wide, meaning any single comparison could still show a real mismatch. A 2022 follow-up by Beyerbach and colleagues, and a 2025 replication by Stadelmaier and colleagues in BMC Medicine, both reached a similar conclusion: trials and cohort studies tend to agree on average.

That hasn’t gone unchallenged. A 2026 re-analysis by Calkins and colleagues argued this close average agreement can arise simply as a statistical consequence of combining many small, uncertain effects, rather than proof that trials and cohorts genuinely agree case by case, and the wide range of individual results supports that caution. This is a live, unresolved disagreement between credible researchers. The fairest summary is that nutrition trials and cohort studies tend to agree on average, while individual comparisons can still diverge considerably.

How do scientists grade the quality of nutrition evidence, and why is that controversial?

The most widely used framework for rating confidence in a body of evidence is called GRADE. In its classic form, GRADE starts randomised trial evidence at a higher certainty than observational, cohort-based evidence, because observational studies aren’t randomised and are more exposed to bias from unmeasured differences between groups.

This matters in nutrition, where large, well-conducted trials of whole diets over many years are often impossible to run. GRADE’s developers have acknowledged this. A 2019 methods paper by Schünemann and colleagues describes an alternative route within GRADE, using a bias-assessment tool called ROBINS-I, that allows well-conducted observational studies to start at high certainty rather than automatically low.

In practice, this hasn’t fully closed the gap. A 2021 survey by Werner and colleagues found GRADE was used to formally rate certainty in fewer than 6% of nutrition systematic reviews in top journals, and when it was used on observational evidence, the certainty rating came out “very low” 61.4% of the time. A 2017 comment paper by Meerpohl and colleagues, all GRADE developers, argued this reflects a genuine shortage of blinded nutrition trials rather than a flaw in GRADE. Researchers behind rival, nutrition-specific scoring tools argue the opposite, saying GRADE’s rules make top certainty ratings for whole-diet questions almost unreachable, whatever the underlying quality of the evidence.

In the UK, the Scientific Advisory Committee on Nutrition (SACN), the government’s independent expert group on nutrition science, updated its evidence framework in 2025 work by Singh and colleagues. It recommends GRADE alongside a related tool for cohort studies, ROBINS-E, while explicitly stating a reservation that GRADE risks undervaluing large, well-conducted cohort studies and unblindable whole-diet trials. Even the UK’s own expert advisers haven’t fully settled this question.

Red meat, two grading systems, two different answers

This disagreement is clearest in the case of red and processed meat, a useful worked example rather than a conclusion.

In 2019, systematic reviews conducted as part of the NutriRECS project, by Zeraatkar, Han and Vernooij and colleagues, all published in Annals of Internal Medicine, used GRADE to assess the evidence linking lower red and processed meat intake to lower rates of death, cancer and cardiometabolic disease. Across large pooled cohorts involving millions of participants, the certainty of this evidence was rated low to very low. The accompanying guideline panel, led by Johnston and colleagues, suggested most adults could reasonably continue their current intake, estimating that reducing intake by three servings a week would prevent around 8 to 9 deaths per 1,000 people over 11 years, a modest absolute effect.

A separate 2020 analysis by Qian, Riddle, Wylie-Rosett and Hu, published in Diabetes Care, applied a different tool, NutriGrade, designed specifically for nutrition evidence, to much of the same underlying research. It rated the certainty of the link between red and processed meat and type 2 diabetes as high, and the link with mortality as moderate. Same broad evidence, two different verdicts.

Meanwhile, the World Cancer Research Fund and American Institute for Cancer Research classify processed meat as a convincing cause of colorectal (bowel) cancer, and the International Agency for Research on Cancer places it in Group 1, the category used for substances with strong evidence of causing cancer in humans. These bodies are answering a related but distinct question, about whether a causal link exists at all, rather than how large the effect is or what a sensible individual recommendation should be, so their conclusions aren’t directly interchangeable with the GRADE or NutriGrade ratings, even though they concern the same food.

I won’t tell you which of these verdicts is correct, because that isn’t settled. What I can tell you is that the disagreement is real, published, and exists among credible, serious research groups rather than between good science and bad science. If a nutrition claim feels contradictory, this is often exactly why.

Are there other ways to test cause and effect in nutrition?

Because no single design answers every nutrition question cleanly, researchers increasingly combine different kinds of evidence.

Triangulation, described by Lawlor, Tilling and Davey Smith in 2016, brings together several methods with different, unrelated weaknesses, on the logic that it’s unlikely entirely different sources of error would all produce the same false result. If several imperfect methods point the same way, that convergence is worth more than any one alone.

Mendelian randomisation uses naturally occurring genetic variation, present from birth and unrelated to lifestyle choices, as a kind of natural experiment to test whether an association is likely causal. A 2022 review by Wade and colleagues found this works reasonably well for individual nutrients with a clear genetic marker, but becomes much harder to apply to complex, whole dietary patterns, where the underlying genetics are far messier.

Target-trial emulation, described by Hernán and Robins in 2016, structures existing observational data as though it were a planned trial, reducing specific, well-understood sources of bias that often creep into simple cohort analyses. None of these methods replaces the others; each adds a different kind of evidence, and each has its own limitations.

Why measurement error and small effect sizes make this harder

Much nutrition cohort research relies on people self-reporting what they eat, often through a food frequency questionnaire, a form asking someone to estimate usual intake over recent weeks or months. This is less accurate than it might seem. A pooled analysis of five validation studies by Freedman and colleagues, published in 2014, compared reported intake against objective biological markers and found food frequency questionnaires under-reported energy intake by an average of about 28%. Error of that size blurs genuine diet-disease relationships, pulling measured effects closer to zero than the true effect might be.

This connects to a long-running argument about how large nutrition’s real effects typically are. Ioannidis has argued, in a 2013 BMJ piece and a 2018 JAMA piece, that many food-disease associations are implausibly large and mostly reflect accumulated bias and selective reporting rather than genuine biology. He points to a study by Schoenfeld and Ioannidis, published in 2013, which looked up 50 randomly selected cookbook ingredients and found most had been linked to either raised or lowered cancer risk in at least one published study, with those effect sizes shrinking substantially once results were pooled properly in meta-analysis.

Other respected researchers see it differently. Satija, Yu, Willett and Hu argued in a 2015 paper that exposures measured carefully and repeatedly, with a plausible biological explanation, tend to produce genuinely reproducible findings, and that the answer to poor measurement is better measurement, not abandoning cohort studies. This, too, is an active, unresolved disagreement between serious scientists.

So, does nutrition science need its own separate evidence hierarchy?

A theme running through a decade of methodology papers is that no single study design, and no single grading system, fits every nutrition question equally well. The growing suggestion, sometimes called fit-for-purpose thinking, is that the right method depends on the specific question, rather than ranking every question against one universal ladder built for drug trials.

Short-term, tightly controlled feeding studies, like Hall’s, suit narrow physiological questions well. Cohort studies remain, realistically, the main option for whole diets and long-term outcomes, because randomising thousands of people to eat differently for decades usually isn’t practical or ethical. Mendelian randomisation and triangulation add value where a genetic angle or a genuinely independent method exists.

Researchers have tried building nutrition-specific alternatives to standard grading. HEALM, developed by Katz and colleagues and published in 2019, is one example. When an independent group, Wingrove and colleagues, applied HEALM to evidence on dietary patterns and all-cause mortality in 2022, the results came out only “moderate” to “insufficient or inconclusive”, suggesting HEALM hasn’t clearly outperformed existing approaches when tested outside the group that built it.

This disagreement about direction of travel is also visible in a pair of perspective papers published side by side in Advances in Nutrition in 2018. Trepanowski and Ioannidis argued nutrition research should lean much harder toward large, well-designed randomised trials wherever practically possible. In the same issue, Satija, Stampfer, Rimm, Willett and Hu argued that genuinely large, simple trials of whole diets are usually unworkable, and that well-designed cohort studies, properly interpreted, remain essential. The gap between them captures the current state of the field about as honestly as anything else here.

No single “best” system has won out. The most honest position, for now, is that nutrition evidence needs to be read study by study and question by question, rather than judged only by where it sits on a single tier ladder that was never built with nutrition in mind.


If there’s one thing worth taking from this, it’s that “the science keeps changing” is usually the wrong way to describe what’s happening. Nutrition evidence comes in many different strengths, researchers openly disagree about how to grade those strengths, and headlines rarely have room to explain either.

This article sets out what current research and methodology papers say about how nutrition evidence is graded in general. It isn’t personalised advice about your own diet or health, and it isn’t a substitute for an individual assessment of your own situation.

If you’re looking to improve your diet and lifestyle, that’s exactly what a registered dietitian specialises in.

If you’ve ever felt talked out of trusting nutrition headlines altogether, save this for the next time a new study makes the news.

This content is for educational and informational purposes only and does not substitute for professional medical advice, diagnosis, or treatment.


References

Beyerbach, J., Stadelmaier, J., Hoffmann, G., Balduzzi, S., Bröckelmann, N. and Schwingshackl, L. (2022) ‘Evaluating concordance of bodies of evidence from randomized controlled trials, dietary intake, and biomarkers of intake in cohort studies: a meta-epidemiological study’, Advances in Nutrition, 13(1), pp. 48-65. doi: 10.1093/advances/nmab095.

Calkins et al. (2026) Matters arising: re-analysis of concordance between randomised trial and cohort evidence in nutrition research, BMC Medicine. [Full volume, issue and DOI details not available in source material; to be verified before further use.]

Estruch, R. et al. (2013, retracted and republished 2018) ‘Primary prevention of cardiovascular disease with a Mediterranean diet supplemented with extra-virgin olive oil or nuts’, New England Journal of Medicine, 378 (republished version). doi: 10.1056/NEJMoa1800389. (Original publication: New England Journal of Medicine, 2013, 368, pp. 1279-1290.)

Freedman, L.S. et al. (2014) ‘Pooled results from 5 validation studies of dietary self-report instruments using recovery biomarkers for energy and protein intake’, American Journal of Epidemiology, 180(2), pp. 172-188.

Han, M.A. et al. (2019) ‘Reduction of red and processed meat intake and cancer mortality and incidence: a systematic review and meta-analysis of cohort studies’, Annals of Internal Medicine, 171(10), pp. 711-720. doi: 10.7326/M19-0699.

Hall, K.D. et al. (2019) ‘Ultra-processed diets cause excess calorie intake and weight gain: an inpatient randomized controlled trial of ad libitum food intake’, Cell Metabolism, 30(1), pp. 67-77.

Hernán, M.A. and Robins, J.M. (2016) ‘Using big data to emulate a target trial when a randomized trial is not available’, American Journal of Epidemiology, 183(8), pp. 758-764.

Howard, B.V. et al. (2006) ‘Low-fat dietary pattern and risk of cardiovascular disease: the Women’s Health Initiative Randomized Controlled Dietary Modification Trial’, JAMA, 295, pp. 655-666.

Ioannidis, J.P.A. (2013) ‘Implausible results in human nutrition research’, BMJ, 347, f6698.

Ioannidis, J.P.A. (2018) ‘The challenge of reforming nutritional epidemiologic research’, JAMA, 320(10), pp. 969-970.

Johnston, B.C. et al. (2019) ‘Unprocessed red meat and processed meat consumption: dietary guideline recommendations from the Nutritional Recommendations (NutriRECS) Consortium’, Annals of Internal Medicine, 171(10), pp. 756-764. doi: 10.7326/M19-1621.

Katz, D.L. et al. (2019) ‘Hierarchies of evidence applied to lifestyle medicine (HEALM): introduction of a strength-of-evidence approach based on a methodological systematic review’, BMC Medical Research Methodology, 19, p. 178.

Lawlor, D.A., Tilling, K. and Davey Smith, G. (2016) ‘Triangulation in aetiological epidemiology’, International Journal of Epidemiology, 45(6), pp. 1866-1886. doi: 10.1093/ije/dyw314.

Ludwig, D.S., Ebbeling, C.B. and Heymsfield, S.B. (2019) ‘Improving the quality of dietary research’, JAMA, 322(16), pp. 1549-1550. doi: 10.1001/jama.2019.11169.

Meerpohl, J.J., Naude, C.E., Garner, P., Mustafa, R.A. and Schünemann, H.J. (2017) ‘Comment on “NutriGrade”‘, Advances in Nutrition, 8(5), pp. 789-790. doi: 10.3945/an.117.016188.

Omenn, G.S. et al. (1996) ‘Effects of a combination of beta carotene and vitamin A on lung cancer and cardiovascular disease’, New England Journal of Medicine, 334, pp. 1150-1155.

Qian, F., Riddle, M.C., Wylie-Rosett, J. and Hu, F.B. (2020) ‘Red and processed meats and health risks: how strong is the evidence?’, Diabetes Care, 43(2), pp. 265-271. doi: 10.2337/dci19-0063.

Satija, A., Stampfer, M.J., Rimm, E.B., Willett, W. and Hu, F.B. (2018) ‘Perspective: are large, simple trials the solution for nutrition research?’, Advances in Nutrition, 9(4), pp. 378-387.

Satija, A., Yu, E., Willett, W.C. and Hu, F.B. (2015) ‘Understanding nutritional epidemiology and its role in policy’, Advances in Nutrition, 6(1), pp. 5-18.

Schoenfeld, J.D. and Ioannidis, J.P.A. (2013) ‘Is everything we eat associated with cancer? A systematic cookbook review’, American Journal of Clinical Nutrition, 97(1), pp. 127-134.

Schünemann, H.J. et al. (2019) ‘GRADE guidelines: 18. How ROBINS-I and other tools to assess risk of bias in nonrandomized studies should be used to rate the certainty of a body of evidence’, Journal of Clinical Epidemiology, 111, pp. 105-114.

Schwingshackl, L. et al. (2021) ‘Evaluating agreement between bodies of evidence from randomised controlled trials and cohort studies in nutrition research’, BMJ, 374, n1864.

Singh, M. et al. (2025) ‘Updating the Scientific Advisory Committee on Nutrition’s Framework for the evaluation of evidence’, British Journal of Nutrition, 134(3), pp. 257-262.

Stadelmaier, J. et al. (2025) [Replication study of RCT-cohort concordance in nutrition research], BMC Medicine, 23, p. 36. [Full article title and DOI not available in source material; to be verified before further use.]

Trepanowski, J.F. and Ioannidis, J.P.A. (2018) ‘Perspective: limiting dependence on nonrandomized studies and improving randomized trials in human nutrition research: why and how’, Advances in Nutrition, 9(4), pp. 367-377.

Vernooij, R.W.M. et al. (2019) ‘Patterns of red and processed meat consumption and risk for cardiometabolic and cancer outcomes: a systematic review and meta-analysis of cohort studies’, Annals of Internal Medicine, 171(10), pp. 732-741. doi: 10.7326/M19-1583.

Wade, K.H. et al. (2022) ‘Applying Mendelian randomization to appraise causality in relationships between nutrition and cancer’, Cancer Causes & Control, 33(5), pp. 631-652.

Wingrove, K. et al. (2022) ‘Using the Hierarchies of Evidence Applied to Lifestyle Medicine (HEALM) approach to assess the strength of evidence on associations between dietary patterns and all-cause mortality’, Nutrients, 14(20), p. 4340.

Werner, S.S. et al. (2021) ‘Use of GRADE in evidence syntheses published in high-impact-factor nutrition journals: a methodological survey’, Journal of Clinical Epidemiology, 135, pp. 54-69.

Zeraatkar, D. et al. (2019) ‘Red and processed meat consumption and risk for all-cause mortality and cardiometabolic outcomes: a systematic review and meta-analysis of cohort studies’, Annals of Internal Medicine, 171(10), pp. 703-710. doi: 10.7326/M19-0655.