Interactive companions

How confident are we that this is a possibly missing link?

Start here — what this is, what it assumes, and how to use it

What this is: an interactive guide to the Bayesian framework of Contextual evidence for categorising interactions and guiding their discovery. The taxonomy assigns each candidate link to one of eight categories from three pieces of evidence — the model prediction Y, the local observation Ol, and the replicate observation Or. That assignment is deterministic, but every source is error-prone, so it carries uncertainty. The framework therefore treats the true category C as unknown and returns a posterior P(C | E) over the eight.

Scope: the three tabs below the setup cards are alternative versions of the framework, one shown at a time. The simple version takes binary evidence, one symmetric error rate per source, a uniform prior, and independent errors. Extension: many replicates relaxes two of those: the replicate evidence keeps its full quantitative count n, and the local axis can err at a different rate in each direction. Extension: informed priors replaces the uniform prior with ecological belief about the species pair, regionally (πr) and at the site (πl). Throughout, the rates start at εY = 0.2, εl = 0.3, εr = 0.1, ρ = 0.15 and f = 0.05, chosen to show the behaviour of the framework rather than to describe any particular system, and each is adjustable with the sliders.

How to use: pick a category on the taxonomy tree. Its error-free signature z_c = (1, 0, 1) becomes the observed evidence E, and the confidence in that label is (1−εY)(1−εl)(1−εr).

Try this: pick possibly missing — predicted and seen elsewhere, but not here. Open Error rates and push εl up: confidence in the label falls and recurrent takes the difference, because the more local sampling misses, the likelier the link was simply overlooked. Then switch to Extension and watch confidence climb as replicates accumulate — and stop at κ.

Mapping observations to truth — understanding the error rates

Each arrow runs from a quantity to the one it is trying to recover, and carries the error rate of that recovery. The green diagonal is what we actually want: how well the model we have, Y, recovers the ground truth. It is never measured directly — but composing the top and right edges gives it. Move the three rates — δ is 1 − accuracy, the number cross-validation reports — and watch the diagonal: it brightens as the model closes on the truth. These sliders are local to this panel and change nothing else on the page.

80.0%apparent accuracy — what CV says
69.5%real accuracy — against the truth
of Y against the ground truth

Pick a link category — click any leaf

Each leaf carries its signature zc = (Ŷ, Ll, Lr): predicted · realised here · realisable in the replicates — the evidence it would produce if no source erred

Category prior (base rates) — uniform by default; edit to break the symmetry

A uniform prior treats all eight categories as equally probable before any evidence arrives, so confidence in the chosen label depends only on the error rates — not on which category it is. Making some categories more probable than others would make the posterior no longer equal to the likelihood. Weights are relative, rescaled to sum to 1 (shown as %).

Error rates εd

how often the prediction would change had the model been fitted to the ground truth — not the error cross-validation reports
the miss rate: the probability that a link realised where we sampled goes unrecorded in Ol
the probability that a realisable link is recorded in none of the replicates

εY and εl drive all versions below. εr belongs to the simple version only: the extension replaces it with ρ and f, and lets the local axis err at a different rate in each direction.

Posterior P(C | E) over the eight categories

possibly missing (no mismatch) 1 mismatched source 2 mismatched 3 mismatched

Key quantities — for the chosen category

Confidence in possibly missing — the posterior probability P(C = possibly missing | E) that the label is correct.
φ — feasibility confidence — a lower bound on the probability that the link is feasible, whichever exact label is right.
κ — maximum contextual confidence — under these symmetric rates, (1−εY)(1−εl): the most that within-system replication alone can deliver.

Likelihood breakdown — every cell updates with the sliders

Categorysignature zcmismatched sources model (Y)local (l)replicate (r) P(E|C)P(C|E)