The Box 2 Bayesian reading of the link taxonomy, in its simplest form (binary evidence, three fixed
error rates, uniform prior, independent errors). Pick a category on the taxonomy tree. Its error-free
signature (Ŷ=1, O_local=0, O_rep=1) becomes the observed evidence, and with a uniform prior
the posterior confidence in that label is (1−ε_Y)(1−ε_local)(1−ε_rep). Drag the sliders — or the point in
the phase space — to see how each error rate moves it, and where the residual mass goes.
With a uniform prior every category is equally likely before seeing evidence, so confidence in the chosen label depends only on the error rates — not on which category it is. Give some categories more prior weight and the headline confidence starts to differ: a rarer category needs stronger evidence to reach the same confidence. Weights are relative and are renormalised to sum to 1 (shown as %).
| Class | signature | model | local | rep | P(E|C) | P(C|E) |
|---|
The full posterior over all eight categories under the selected consistent-evidence scenario, so the eight lines sum to 1 at every R. Each category has its own colour; the chosen category and its replicate counterpart are drawn bold with markers. Switch the scenario to see the mirror image — and note each plateau equals the bar-chart posterior in the ε_rep→0 limit, so this is the bar chart emerging as replicates accumulate. The dashed ceiling = (1−ε_Y)(1−ε_local) is the most confidence replicate evidence can ever buy: no amount of replicate data resolves the model (ε_Y) or local (ε_local) errors, so the chosen category can climb to that line but no further.
The same model as above, inverted: instead of "how confident am I after R replicates?", ask "how many replicates of consistent evidence do I need to reach a target?" The two lines are the chosen category and its replicate counterpart — the two halves of one confusion cell (same Ŷ, O_local; opposite O_rep). One is corroborated by presence (it keeps being seen elsewhere), the other by absence (it keeps not being seen), so this contrasts how fast a link is confirmed by detections vs by absences. Pick any category on the tree to re-pair; ρ and f are shared with the panel above.
Both categories' confidence rises with consistent replicates toward the shared ceiling = (1−ε_Y)(1−ε_local) — here the y axis is scaled so that ceiling is the top of the plot, so a line reaching the top edge has extracted everything replicates can give. The y values are real confidences, and so is the target: the orange line sits at the absolute confidence you asked for, and each R* dot marks where a curve crosses it. The target slider stops at the ceiling — you cannot ask for more than replication can buy, so tightening ε_Y or ε_local is what raises the bar. The presence-corroborated line is steeper (fewer replicates); the absence-corroborated one lags. That gap is the asymmetry: a detection is a rarer, more informative event than an absence, so presence resolves faster for the same ρ and f. Both R* values diverge at the crossover f* = ρ(1−ε_local). This panel shares ρ and f with the accumulation panel above.