How we know the numbers are right.

Every figure on BedsideDx comes from the published medical literature, and we check each one against the original study before it goes on the site. Here is how we find the evidence, judge it, and report it — and how you can verify any number yourself.

Every number is real, and verified

We do not generate diagnostic statistics — we report them from peer-reviewed research, and we confirm each one against the source before publishing it. The study’s authors, year, journal, and the actual numbers all have to match. Every figure on the site links to its citation so you can check it yourself, and if a study’s claim doesn’t hold up on inspection, it doesn’t go on the site. We never estimate a value to fill a gap.

We grade the evidence, not just report it

A striking number from a single small study is not the same as one pooled from every study ever done on the question. So alongside each finding we show how strong the evidence behind it actually is:

  • Strong — drawn from the most rigorous evidence available: a systematic review or meta-analysis that pools this specific finding’s accuracy for this diagnosis against an independent reference standard. What earns the tier is the study’s design — never the journal it appeared in.
  • Limited — based on a single well-conducted study, or one study drawn from a larger review. Real evidence, but a narrower foundation.
  • Insufficient Evidence — we searched and found nothing that met our bar. Instead of inventing a number, we label the finding honestly and note what is missing. You still see the finding and its clinical context — just not a statistic we can’t stand behind.

The grade reflects the quality of the evidence, not the size of the number. A large likelihood ratio from one small study still rests on limited evidence, and we label it that way.

How we find the evidence, and what clears the bar

For every finding we work down a hierarchy:

  1. Rigorous syntheses first. Systematic reviews and meta-analyses conducted to recognized standards for searching the literature and appraising risk of bias. When one has already answered the question, we use its pooled results as published.
  2. Then well-designed primary studies. If no such synthesis exists, a single study can qualify — but only if it was done well: an appropriate patient population, a blinded comparison against a true diagnostic standard, an adequate sample, and complete reporting of its accuracy.
  3. Otherwise, we say so. If nothing meets that bar, the finding is marked Insufficient Evidence.

We are deliberately strict about two things. The study has to be about the same diagnosis we are applying it to — we don’t borrow a number from a related condition or stretch one study’s result across several diagnoses. And we don’t anchor a rating on preprints, conference abstracts, or exam textbooks; those can inform context, but a rating has to rest on peer-reviewed evidence.

What the ratings mean for you

We lead with the likelihood ratio because it tells you how much a result should actually move you for the patient in front of you: a large ratio makes a diagnosis much more likely when the finding is present, and a very small one makes it much less likely when the finding is absent. Unlike sensitivity and specificity, it applies directly at the bedside.

So you don’t have to do the math in your head, we translate every ratio into a plain rating using the long-established Sackett thresholds:

  • Very helpful — a large, often decisive shift
  • Helpful — a moderate, frequently useful shift
  • Somewhat helpful — a small shift that occasionally matters
  • Minimally helpful — a shift too small to act on by itself
  • Not helpful — no meaningful shift
  • Inconclusive — the study’s headline number looks useful, but the range of values it is actually compatible with runs from no effect at all up to a large one. The evidence cannot settle the question either way, and we say so rather than pretending otherwise in either direction. The number itself stays on the page, next to the rating we are withholding, so you can see exactly what the study reported.

Where a rating needs a qualification, a short note sits directly beneath it naming the problem — CI is the confidence interval, the range of values the study is actually compatible with. Hover over the note for the full sentence. There are four, and a finding carries at most one:

  • No CI — the paper published its result but no range around it, and not enough detail for anyone (its authors, us, or you) to work one out. The rating reflects the single number the study reported and nothing more.
  • Wide CI — the study did report a range, and that range runs from no useful effect up past a clinically important one. On an Inconclusive finding it means the range is too wide to rate the number at all; on a Not helpful finding it means the central result was flat but the study was not large enough to rule a real effect out.
  • No CI, small study — no range, from a single study of fewer than fifty patients. Here we withhold the rating altogether: the finding reads Inconclusive and keeps both its number and its evidence grade. One number from a handful of patients, with no measure of its uncertainty, is not a rating we can stand behind — but a study was found and named, so this is never Insufficient Evidence. What we withhold is the rating, not the evidence.
  • Derived CI — the source published no range for the likelihood ratio, but it reported enough for one to be calculated from its own counts, so the range you see is ours rather than the paper’s, and we say so instead of letting our arithmetic pass for theirs. We are working through the findings that allow this, so this note will appear on more of them over time.

How much a finding shifts probability, and how sure anyone can be of that shift, are two different questions, and we answer them separately. The rating comes from the study’s central result. Beside it we look at the range of values the study is compatible with — its confidence interval — and where a paper reports enough for that range to be calculated, we are working it out ourselves and marking it Derived CI; until a finding’s range is in hand it is marked No CI. When the range includes the possibility of no effect at all, the finding does not get full credit. It is called Inconclusive, unless the whole range is narrow enough to be clinically unimportant anyway — roughly 0.5 to 2.0 on the likelihood-ratio scale, the band Jaeschke, Guyatt and Sackett described as altering probability to a small and rarely important degree — or the central result itself sits in that flat zone. In those two cases, Not helpful is a verdict the evidence has actually earned, and we give it. And where no range can be obtained at all, we say so plainly rather than quietly granting the finding the same standing as one whose precision was reported and held up. We used to rate every interval that touched no effect as Not helpful outright; that penalized the studies candid enough to publish their uncertainty and flattered the ones that published none.

Some findings are the reverse of what you’d expect — a positive result argues against the diagnosis. The classic example is reproducible chest-wall tenderness in suspected acute coronary syndrome, where finding it lowers the probability of a cardiac cause. We label these by their true clinical meaning so there’s no ambiguity.

The probability calculator

For subscribers, the calculator does the arithmetic. Enter your starting (pre-test) probability and the result you observed, and it returns the updated (post-test) probability using Bayes’ rule — the standard way to combine a prior probability with a test result. It computes wherever we published a rating, and declines where we did not: it will not multiply by a number the study could not settle, or by one from a small lone study with no measure of its uncertainty. Where it does compute but the result needs a caveat — no published range behind the ratio, a wide one, or a finding the evidence showed barely moves the probability — it says so on a line beneath the answer rather than leaving you to guess.

What we are upfront about

Diagnostic evidence is never perfect, and we would rather you know the caveats than discover them later:

  • The studied patients may not be yours. Accuracy studies are often done in particular settings — a referral clinic, an emergency department — and may not translate to a different population. Where the mismatch is large, we flag it.
  • Some findings are hard to reproduce. A number of physical signs are elicited or interpreted differently from one examiner to the next, which limits even a strong result in everyday practice.
  • Yardsticks change. What counted as the diagnostic gold standard decades ago has sometimes been surpassed — ultrasound replacing the chest X-ray for some questions, for instance — so older numbers can understate how an exam compares today.
  • Flattering results get published more. Studies showing a test works are likelier to appear in print than those showing it doesn’t, which can make a sign look better than it really is. We favor the most comprehensive syntheses available, but no source removes this risk entirely.

← Back to about