Frequently Asked Questions

The Problem
A p-value of 0.04 and a p-value of 0.01 are both reported as "statistically significant" — but they can represent vastly different levels of evidence quality that p-values alone do not reveal.

The Solution
Complete statistical evidence requires three dimensions:
- Significance (p): Compatibility with null hypothesis
- Fragility (fr): Classification stability
- Robustness (nb): Distance from neutrality

What is the p-fr-nb triplet?
Complete statistical evidence requires three numbers:
- p: p-value (significance)
- fr: fragility quotient (stability)
- nb: neutrality boundary metric (robustness)
Reporting only p-values provides partial evidence. Complete statistical evidence requires all three dimensions. Clinical decisions also need the absolute effect size: the triplet plus effect size is complete evidence.

Why not just use p-values?
P-values answer one question: "How compatible are the data with no effect?"
They don't tell you:
- How stable is this classification? (fragility)
- How far from neutrality is the result? (robustness)
These are complementary dimensions that p-values do not measure.

What is Pattern (1,1,0)?
Pattern (1,1,0) = thin evidence, consistent with a trivial effect:
- 1: p-significant
- 1: fr-unstable
- 0: nb-near (close to neutrality)
In a study of 129 published two-arm, binary-outcome trials, 24 of the 77 significant trials (31.2%) showed this pattern, defined there as p ≤ 0.05, MFQ ≤ 0.10, and RQ < 0.075. Simulated trials showed it at similar rates (24–32%) whether the true effect was strong, moderate, or absent, so the pattern marks a precarious boundary zone rather than identifying false positives (Heston, Cureus, 2025). The p-value says "significant." Complete statistical evidence says "remain skeptical": treat the effect as clinically negligible unless it is replicated with an nb farther from neutrality.

How is this different from clincalc.com?
The clincalc.com fragility calculator implements a different calculation for the FI than FragilityMetrics.org. We implement the Heston Fragility Index, a modification of the original Walsh (2014) method.

The original Walsh et al. (2014) method converts non-events to events in the arm with fewer events until significance is lost, and it does not say what to do when event counts are tied. The Heston FI adds an explicit tie rule, if events are tied, toggle the smaller arm; it is also bidirectional, defined for nonsignificant results, and defined as a true minimum.

Our testing shows that Clincalc.com appears to do the following: when events are equal between arms, its calculator defaults to toggling the control group regardless of arm size. This can produce higher FI values when the control arm is larger than the experimental arm.

Example: Control (10 events, 22 non-events, 32 total patients) vs Experimental (10 events, four non-events, 14 total patients). Per the Heston FI tie rule: events tied → toggle the smaller arm (experimental group). This implementation is used on FragilityMetrics.org (result: FI=1). In this case, Clincalc.com defaults to toggling the control arm when events are tied (result: FI=2). These are the results from 12/26/2025 (FragilityMetrics, ClinCalc).

Additional features of FragilityMetrics.org:
- We report MFQ (Modified-Arm Fragility Quotient), which is allocation-fair: it divides FI by the size of the arm actually toggled, minimizing distortion from unequal allocation
- We provide the Global Fragility Index (the minimum number of patients moved between any cells of the table, total N fixed, to flip significance) and the Global Fragility Quotient (GFQ = GFI/N), the recommended fragility metric for 2×2 tables
- We provide the complete p-fr-nb triplet, not fragility alone
- We report RQ (Risk Quotient), the robustness (nb) metric, and NDI (Neutrality Distance Index), its count form; clincalc.com calculates neither
- We report the absolute effect size (risk difference with 95% CI, and NNT) alongside the relative risk
- The FI is bidirectional: if the baseline p-value is nonsignificant, the FI is the number of toggles required to cross the p = 0.05 threshold and become significant, so no separate "reverse" FI is needed.
- Open source code available for verification
- We encourage testing of our software and have made it open source (CC-BY 4.0). Please report any bugs on the GitHub discussion board.

Is this calculator free?
Yes, completely free. No registration required. Open source implementation.

What outcome types are supported?
- Binary (2×2) on this site: GFQ + RQ, with FI, FQ, MFQ, GFI, NDI, and the absolute effect size also reported (2×2 calculator)
- Survival on this site (hazard ratio with confidence interval): SFQ + SRQ (survival calculator)
- See the Fragility Metrics Toolkit for Python notebooks covering continuous, multi-group (ANOVA), r×c, fixed-margin 2×2, diagnostic, ordinal, survival, correlation, and single-arm benchmark outcomes.