| Field | Value |
|---|---|
| Engine | Qwen/DeepSeek-V4-Flash Statistics auditor |
| Rules | ReproAI/REPRO_STANDARDS.md |
| Audit date | 2026-09-16 |
| Package | FragilityPrinciple v1.0 — Zenodo doi:10.5281/zenodo.21834953 (MSID RSOS-261834) |
| Mode | Live (ran author code end-to-end) + static + independent reimplementation |
| Host env | win32 · Python 3.8.10 (numpy 1.24.4, scipy 1.10.1, pandas 2.0.3, matplotlib 3.7.5) · R 4.6.1 (Rscript at C:\Program Files\R\R-4.6.1\bin) |
| Host toolchain | R and Python are both installed and used on this system. The author code is Python and was re-run live; the independent second-language recompute was done in base R (no author modules). Other statistical packages (Stata, SPSS, MATLAB) are not installed here; in a companion recheck they would need to be translated to R or Python (both available) rather than run natively. |
| Version note | Package declares Python≥3.10 / numpy≥1.26. Re-ran on older ABI-compatible stack with fixed seeds; all values reproduce at reported precision (only ~1e-15 float/BLAS drift). |
| Scope | fragility-principle-v1.0 (code 15 files, results 14 files). No git history (Zenodo snapshot). |
| Severity | Count |
|---|---|
| P0 blocker | 0 |
| P1 high | 0 |
| P2 medium | 5 |
| P3 low | 3 |
Verdict. This is a genuinely reproducible package of unusually high quality. Every
table (1–4), the abstract statistics, the §3.2 critical ratios, and both §3.3 EEG
confirmatory arms regenerate exactly at reported precision from the shipped seeds,
and the confirmatory classification is independently confirmed from the shipped raw
file with hand-written logic. There are no wrong numbers: no P0/P1. The findings are
all traceability/provenance and scope issues — the strongest being that three
quantitative prose claims (6000-family geometry check, null VR≈0.66 calibration, the
Metric-T "R exceeds |T| threefold" claim) are not backed by any shipped script or
data, and the promised Metric-T DC(M)/T/R values are never reported in §3.4.
| Script (mine / shipped) | Re-implements | Manuscript exhibit | Status |
|---|---|---|---|
bia_sobol.py (7) |
Two-layer functional ANOVA + VR | Table 1, Fig 3–4 | ✅ exact |
bia_sobol3.py (7) |
Three-layer functional ANOVA, PS scale | Table 2, ESM S3 | ✅ exact |
bia_confirmatory.py (7) |
Hedges'-g response fit + cluster bootstrap | Table 3, §3.2 | ✅ exact |
okimo_table6.py (7) |
Empirical 672-spec grid, FANOVA, VR | Table 4 | ✅ values match; ⚠ schema/provenance drift |
eeg_confirmatory.py (8, 0x7) |
EEG reconstruction arms | §3.3 | ✅ exact narrative |
fragile_geometry.py (7) |
Geometry + weighted-Φ range | §2.4 | ✅ kappa=2.5 balance confirmed |
sobol3.py --self (7) |
ANOVA exactness self-tests | §2.5 | ✅ pass |
independent_recompute.py (mine, numpy-only) |
Re-derive Table 3 coefficients + r* from raw CSV | §3.2 | ✅ agree (<3e-5; critical ratios exact) |
independent_recompute.R (mine, base R) |
Re-derive Table 3 coefficients, 2000-resample cluster bootstrap CIs, verdicts, classification + r* from raw CSV | §3.2/Table 3 | ✅ 7/7 classifications OK; all verdicts & r* exact (diff=0.0000) |
delay_transition.py (8) |
ESM §S1 delay sweep | ESM S1 | ⚠ analysis not runnable (raw not shipped); derived ✓ |
make_figures.py (7) |
Figures 1–4 | Figs 1–4 | ✅ all four build from stored CSVs |
Legend: ✅ identical at reported precision · ≈ close/method-explicable · ⚠ discrepant · ➖ not independently checkable.
Manuscript vs shipped/recomputed (medians): null S_E 0.570/0.570, ST_E 0.740/0.740, VR 0.046/0.046; effect-only 0.744/0.069; moderate 0.882/0.364; strong S_E 0.702/0.702, ST_E 0.908/0.908, VR 0.984/0.984. Percentiles agree. Abstract "0.74→0.91", "VR 0.046→0.984" ✅.
null ST_E 0.598/0.598, ST_I 0.138/0.138; strong ST_E 0.886/0.886, ST_I 0.008/0.008. §3.1 "inference total effect 0.138→0.008; estimator 0.598→0.886" ✅. Per-replicate component sums = 1.0 (asserted) ✅.
Xc a 0.137[0.128,0.145] nuisance-sensitive; PhA 0.159 nuisance-robust; Xc_H 0.130[0.122,0.138] nuisance-sensitive; R_H/Z_H/BIVA non-target; XcZ nuisance-robust. §3.2 critical ratios r_BMI Xc +1.002[0.922,1.091], Xc_H +0.981[0.902,1.068]; r_H Xc_H +1.698[1.482,1.952] — all regenerate exactly; independent numpy recompute from raw CSV agrees (<3e-5 for coefficients, 0.0000 for critical ratios). Independently confirmed in base R (R 4.6.1): a from-scratch read of bia_confirmatory_raw.csv + 2000-resample cluster bootstrap (seed 20260806) reproduces all seven index classifications (7/7), every verdict, and every critical ratio to 4 decimals (diff = 0.0000). ✅
T7-Pz S_E 0.410, S_EP 0.279, ST_I 0.158, VR 0.124, p_VR 0.695, D 0, Φ undefined, Φ⁰ 0 = shipped. AF3-T7 similarly ✅. But current okimo_table6.py no longer produces the p_VR column (it writes an 18-column schema); the p_VR values trace to okimo_fragility.py analyse() (seed 20260807). No single shipped script regenerates the full shipped row set.
"none of 672 specifications reached significance" — min adjusted p = 0.0603 (T7-Pz) and 0.0503 (AF3-T7), both > 0.05, D=0 ✅. §3.3 common-source arm: six indeterminate + ciPLV equivalent target loadings, all nuisance material ✅; cross-mix arm: 6 material target loadings, empty R_robust ✅ ("no common consensus interval").
The diagnostics rely on per-specification sign of the directed effect. A common
bug class is taking the sign of a two-sided statistic or of |T|; this package does
not do that.
Code path (fragility_sim.py:237-299):
- sign = int(np.sign(g)) where g = (x_B.mean() − x_A.mean())/s_p — the directed
Hedges' g (group B = AB, group A = CB), not |g|, not the p-value.
- Mann–Whitney decision_sign uses the rank direction sign( U/(n_A n_B) − 0.5 )
(probability of superiority), not the (two-sided) p.
- Φ⁰ (phi_zero, line 313) counts r.sign — the directed-effect sign — for all
specs; Φ (fragility_indices) counts decision_sign among significant specs.
Empirical distribution of signed effects (sign_check.py, confirmed against an
independent per-cell recomputation):
| Pair | g range | positive (AB>CB) | negative | zero | mean g | Φ⁰ |
|---|---|---|---|---|---|---|
| T7-Pz | +0.050 … +0.954 | 672 | 0 | 0 | +0.4365 | 0 |
| AF3-T7 | +0.031 … +0.785 | 672 | 0 | 0 | +0.4207 | 0 |
Every one of the 672 directed effects is positive; none is negative or zero.
- r.sign == sign(g) for all 672 specs (True) — Φ⁰ uses the correct directed sign.
- r.decision_sign == independent recomputation (True) — Φ's sign is also correct.
- 30 (T7-Pz) / 22 (AF3-T7) Mann–Whitney specs have a rank-based direction that
opposes the mean direction (decision_sign ≠ sign(g)) — confined to MW rows. This
is intended (rank vs mean functionals) and does not affect Φ⁰ (Φ is undefined because
D=0).
Why Φ⁰ = 0 is real and does not contradict Metric‑T [18].
Φ⁰ = 2·min(%pos, %neg) = 2·min(1.0, 0.0) = 0 because all 672 point estimates share
one sign. That is a genuine property of this fixed feature (each electrode pair), not
an artifact of taking sign(|T|): it means on the AB-vs-CB contrast at the group-mean
level, no specification flips direction. It does not contradict [18] because the two
diagnostics average over different axes:
1. Φ fixes the feature and varies the analyst path (7 estimators × 16 preprocessing
× 6 inference). Metric‑T fixes the analytic path and varies the feature/connection.
A single feature can be direction-consistent across specs (Φ⁰ = 0) while different
features or a metric pair reverse — i.e. DC(M₁) − DC(M₂) can be non-zero even
when every spec on this one feature is same-signed. The axes of variation are
orthogonal, so agreement on one is silent about the other.
2. Differentiability: Φ⁰ answers "do the point-estimate directions agree within the
spec space?"; Metric‑T answers "does a metric's direction-consistency differ from
another metric's across features?" A common positive bias on the AB cohort would
push every estimator's mean up together (all g > 0, Φ⁰ = 0) while leaving metric-to-
metric feature reversals (T) intact — exactly the "nuisance that biases every
specification in the same direction remains invisible" caveat the manuscript itself
states (§3.5).
So Φ⁰ = 0 is "all point estimates same-signed" directional agreement, which is a
statement about within-spec-space consensus on one feature, not a statement about
across-feature/across-metric reversal that [18] quantifies; it neither corroborates nor
refutes Metric‑T's T (which the manuscript, as noted in finding ADV‑004, never actually
reports).
Detailed entries, evidence (file:line/path) and fixes are in artifacts/risk_register.json.
okimo_table6.py does not regenerate shipped okimo_table6.csv (schema drift; p_VR column no longer produced by that script; comes from okimo_fragility.py analyse).delay_transition.py analysis stage fails: delay_transition_raw.csv not shipped; only derived delay_transition.csv present (which does support the ESM §S1 sign-change claim).DC(M), T, R never reported; the "threefold" claim is unsubstantiated.okimo_table6.py call it "Table 6"; MS Table 4) and institution spelling (PAZ vs Paz).sobol3.py prints Japanese self-test labels; fails on a cp1252 console without PYTHONIOENCODING=utf-8 (self-tests themselves pass). Portability, not numeric.Legend for symbols used throughout: ✅ = identical at reported precision; ⚠ = discrepant/provenance; ➖ = not independently checkable (with reason); ≈ = close/explicable.
[RECOVERED]/[SPEC]/[RECONSTRUCTED] labelling is applied throughout the code, and the manuscript explicitly withdraws the historical EEG boundary, treats the BIA run as a "definitive rerun of the reconstructed model", and distinguishes pilot vs confirmatory rules. This is a model for the field.audit_reproducibility.py, 58 checks, exit 0 = pass; passed here), and make_figures.py reads values from CSVs rather than embedding transcribed literals..npz files contain derived connectivity only, pseudonymous subjects n1–n14 (n7 absent per a priori exclusion, CB 4 / AB 9), group labels only — confirmed by inspection.incomplete flag.okimo_fragility.py extract_preprocessed cannot be run from the package. The empirical pipeline was validated only against the shipped derived .npz matrices, not from raw recordings.bia_model.py, MVAR leak operators in fragility_sim.py) are calibrated to surviving outputs/spec and labeled RECONSTRUCTED; their parameter values (PHI0, TBW exponents, SIG_LNR, etc.) are not independently verifiable against a lost original.p_VR but permutation CIs about any single seed were not re-run.exec_check/)console_log.txt — console evidence with ==== START/END ==== markers (audit, okimo_table6, bia_confirmatory, bia_sobol3, bia_sobol, eeg_confirmatory ⨯2, sobol3, fragile_geometry, make_figures, delay_transition).independent_recompute.py — numpy-only recompute of Table 3 from bia_confirmatory_raw.csv.sign_check.py — verification of sign extraction + signed-effect distribution across the 672-spec grid (both pairs), independent rank-direction recomputation.independent_recompute.R — base-R recompute of Table 3 (OLS + 2000-resample cluster bootstrap, seed 20260806) from bia_confirmatory_raw.csv; 7/7 classifications and all r* match.artifacts/architecture_report.json, artifacts/risk_register.json, artifacts/advisory_plan.json.fragility-principle/figures/fig{1..4}.{pdf,eps,png} — regenerated.extracted/ — untouched reference (SHA-256 check-summed); exec_check/ — working copy.==== START (status: OK) engine=Qwen/DeepSeek-V4-Flash rules=REPRO_STANDARDS.md date=2026-09-16 host=win32 py=3.8.10 numpy=1.24.4 scipy=1.10.1 pandas=2.0.3 matplotlib=3.7.5 package=v1.0 mode=live ====
[audit_reproducibility.py] 58 checks; 0 failed
[sobol3 self-test] S sum=1.000000000000; VR direct matches; I-constant S_I~1e-34 (pass; utf-8 console)
[fragile_geometry] R_robust=(-0.1949, +0.1172) numeric max-abs-err=3.44e-05; kappa=2.5 -> Phi_max=1.0 balance=True
[bia_confirmatory.py] all 7 index classifications + r* match Table 3; EXIT=0
[independent_recompute.R] R 4.6.1: 7/7 classifications OK; all verdicts & r* exact (diff=0.0000); EXIT=0
[bia_sobol3.py] Table 2 medians match shipped exactly; EXIT=0
[bia_sobol.py] Table 1 medians/percentiles match shipped; EXIT=0
[okimo_table6.py] Table 4 values match; schema differs (p_VR absent); EXIT=0
[eeg_confirmatory] common arm: 6 indeterminate + ciPLV equiv (matches); crossmix arm: 6 material, R_robust empty (matches)
[make_figures.py] fig1-4 pdf/eps/png written; EXIT=0
[delay_transition.py] analysis FAILED (delay_transition_raw.csv not shipped) -- expected per ADV-002
==== END (status: OK) ====
Generated 2026-09-16 by ReproAI audit (REPRO_STANDARDS.md). Source: REPRO_REPORT.md