Stateam logoSTATEAMScience Technology Analytics Team

Verified clinical analytics · Legacy modernization · Available

Parity

Move off legacy without losing the evidence.

Parity logo

Overview

Parity is the evidence layer for clinical and legacy analytics. It converts legacy statistical programs and COBOL to R and Python, proves every regenerated output against the original cell for cell, and keeps a hash-verified audit trail behind every number.

Anyone can translate code. The hard part is proving the new code produces the same numbers, and being able to show that to a reviewer. Parity does both, automatically, behind one quality gate that does not negotiate.

At a glance

Built forBiometrics teams, contract research organisations, sponsors and government agencies
ConvertsLegacy statistical programs and COBOL
TargetsR and Python, both, as two independent implementations
Acceptance barZero tolerance by default, cell by cell
RunsFully offline, including air-gapped networks
RequiresPython 3.9 or later for the core path

The problem

Migrations don't stall on translation. They stall on proof.

Translation is the easy half

Converting legacy code to R or Python is a solved problem. Demonstrating that the regenerated table matches the original to the last decimal, and showing a reviewer that evidence, is not.

The differences are silent

Rounding rules, missing-value ordering and row-lag behaviour differ between platforms. Each produces a plausible wrong answer, never an error.

Verification eats the senior people

Double programming and cell-by-cell review is manual, unbudgeted, and consumes exactly the statisticians you cannot spare.

What a reviewer asks, and Parity answers

  • Show me the old output and the new one, side by side.
  • Which cells differ, and by how much?
  • Which package versions produced this?
  • Has anything been edited since it was produced?
  • Who checked it, and against what?

Five modules, one set of rules

You do not learn the module names. You say what you want, such as "convert this program to R and Python", "analyse this data" or "QC this before we submit", and Parity picks the module, runs it, and reports where the output went. Every module shares one output layout, one audit trail, one environment record and one quality gate.

Parity Convert

Legacy statistical programs and COBOL to R and Python, proven cell for cell against the original output.

Parity Analyze

Profiles any dataset, fixes the analysis plan before the comparison, and runs it reproducibly.

Parity Lock

Pins every Python and R package per project, with a Dockerfile and a version ledger.

Parity Evidence

Systematic literature search with the full PRISMA record, and meta-analysis.

Parity Gate

A deterministic quality gate, plus an independent review of the outputs.

How a conversion runs

1. Inventory

Every step, macro, dataset, title and footnote, plus a risk register naming the constructs in this program that behave differently on the new platform, with line numbers.

2. Specify

Input, filters, derivations in order, statistics and options, rounding rule and output layout, written down before any code.

3. Implement

R and Python, both. Two independent implementations agreeing is itself evidence.

4. Verify

Each regenerated table is compared with the original, cell by cell, at zero tolerance by default.

5. Attest

Independent re-derivation from the original source, then the gate. The traceability matrix is closed.

A conversion is only signed off when there are no blocking differences. The gate reads the result, so nothing is signed off on optimism.

What "verified" means

Clinical cells are rarely bare numbers. A cell such as "142 (78.5%)" is split into its literal text and its numbers: the text must match exactly and the numbers to a declared tolerance. The default tolerance is zero, and raising it is a decision that is written down and justified, never a quiet setting change.

ClassWhat it meansSeverity
ValueA number differs beyond the declared tolerance, or one side is missingBlocking
TextWording or layout differs: a label, a footnote, a spanning headerBlocking
StructureRow or column counts differ, or a keyed row cannot be pairedBlocking
FormatSame value, different rendering, such as 0.5 against 0.50Advisory

The differences that don't announce themselves

Parity ships a risk register that finds these in your source, by line number, before anyone writes a line of R or Python, together with the correct equivalent for each.

ConstructWhat goes wrong on the new platform
RoundingMany legacy platforms round half away from zero; R and Python round half to even, so 2.5 becomes 3 in one and 2 in the other
Missing valuesWhere a missing number sorts below every value, a filter such as x < 5 keeps missing rows; elsewhere it drops them and the population count changes
Lagged valuesA lag inside a condition is not always the previous row; in R and Python it always is
MergingA legacy merge is not a database join; many-to-many merges behave differently
DenominatorsClinical percentages usually use the population count, not the count of non-missing values

Legacy modernization, including COBOL

The same problem in COBOL has harder consequences, because the silent differences are money rather than statistics.

Fixed-point arithmetic is exact

Packed decimal fields have no representation error. Ported to a floating-point type, money changes at the cent, on every record, silently.

Truncation is the default

COBOL discards excess digits rather than rounding, and when it does round it rounds half away from zero, not the way Python and R do.

Source order is behaviour

With PERFORM THRU, every paragraph between the endpoints runs by fall-through. Reordering paragraphs changes the result.

The comparison engine reads fixed-width report output directly, and the run record, tamper evidence, environment lock and quality gate carry over unchanged. For programs whose output is money, Python is the defensible target; this is stated in the documentation, not discovered in a comparison.

Analysis done in the right order

Profile first, plan second, and only then analyse. Before anyone runs a model, the profile catches disguised missing codes such as -99 in a score column, categories that differ only by capitals or spaces, site IDs that lost their leading zeros, and two columns that are really the same variable.

  • The estimand, stated before the comparison runs
  • The estimate and its confidence interval first, the p-value after it
  • Assumption checks reported whether or not they passed
  • Everyone excluded, with reasons
  • A saved script that reproduces every number

Every number carries its own evidence

Append-only

Nothing rewrites a previous run. A correction is a new run that references the one it supersedes.

Tamper-evident

Output fingerprints (SHA-256) are re-checked by the gate. A file edited after the run is detected and blocks release.

Self-contained

Inputs are copied into the run, so it stays verifiable after it is archived.

Portable

Paths are stored relative to the run, so an archived folder still verifies after it is moved.

Reproducible environments

A number is only reproducible if the software stack is recorded. For every project, Parity Lock records every Python and R package and version, the interpreter builds, operating system and code version; writes installable pin lists for both languages and a Dockerfile that rebuilds the whole stack; and keeps an append-only ledger of every environment the project has used. One command verifies a lock and lists every package that has drifted.

Evidence synthesis

Literature search

PubMed, Europe PMC and ClinicalTrials.gov searched directly; subscription databases ingested from your exports. Records are de-duplicated, trial registry records stay visible so publication bias can be assessed, and exact queries, dates and hit counts are captured for the PRISMA report.

Meta-analysis

Standard effect measures and estimators, heterogeneity reported with its uncertainty, prediction intervals by default, and small-study tests with their limitations stated. An independent script redoes the analysis so agreement is real evidence.

Checked against published values

Parity's statistics are tested against reference results. On the standard BCG vaccine dataset (13 trials, risk ratio, random effects), Parity agrees with the established R package metafor to four decimal places.

QuantityParitymetafor
Random-effects estimate (log RR)-0.7145-0.7145
Standard error0.17980.1798
95% confidence interval-1.0669 to -0.3622-1.0669 to -0.3622
tau-squared0.31320.3132
I-squared92.22%92.22%

Built to be validated

Installation qualification (IQ)

Package and file checksums of the deployed software, a platform record and an environment lock: the machine-readable half of the IQ.

Operational qualification (OQ)

Run the automated verification suite and keep its output and exit status as evidence for the named requirements.

Performance qualification (PQ)

A parallel run: convert outputs you have already signed off, and evidence zero blocking findings at zero tolerance.

The control that matters: no statistic, comparison or verdict is produced by a language model. The model decides what to run and explains the result; a deterministic script produces the number. That boundary is what makes Parity usable for regulated work, and it is testable.

Deploy it the way your organisation works

Analyst plugin

Individual analysts drive it directly. The simplest start.

Shared output root

One auditable archive for a team, with one backup policy.

Standalone scripts

For CI pipelines and scheduled jobs, the deterministic half on its own.

Containerised

For regulated work and multi-year reproducibility.

Only the literature module reaches out to the internet, to three public services. Everything else runs fully offline, including on air-gapped networks.

What Parity does not claim

A passing gate is not a correct conclusion

It means the run is intact and self-consistent. Whether the analysis was the right one is a judgement, and a person signs it.

Screening stays a human decision

The screening assistant suggests, with reasons, and includes anything uncertain. It does not decide.

Run records support compliance, they are not a compliance system

Where 21 CFR Part 11 applies, run folders are records to be held inside a validated document system, not a substitute for one.

Bitwise identity has a limit

Maths libraries differ between machines in the last decimals of model fitting. Where that matters, the container is the unit of reproducibility.

Getting started: one table

1. Pick one signed-off table

A demographics or disposition table you have already delivered, with its output.

2. We convert and verify

R and Python implementations, compared cell by cell at zero tolerance, with the risk register and traceability matrix.

3. You review the evidence

The comparison report, the traceability matrix and the gate verdict: the same artefacts a reviewer would ask for.

4. Decide on the programme

Scope the rest against real evidence rather than a proposal.

Every engagement ships with a user guide, technical documentation and a deployment guide that includes the IQ/OQ/PQ mapping. To arrange a pilot or a capability briefing, contact us.

Parity news