Capabilities

A unique combination of modular, purpose-built capabilities​

Decentralised, at source​

data interface​

All required data​
pre-processing 1​

Full data usage
​
control, logs & audits​

Reversible​

pseudonymisation​

Irreversible​

anonymisation​

Configurable privacy treatment
for optimal utility 2

Pseudonymised/identifiable​
linkage at record level​

Anonymised​

linkage at record-level​

Computation at source
for federated analytics 3

1 Including data quality checks of completeness and validity, structuring of unstructured data, and data model standardisation​
2 Optimising utility for any required level of privacy protection​
3 Statistical attributes of interest in each source for confirmed overlapping or non-overlapping cohorts of interest​​

Data Controllers fully control FITFILE-deployed software, and FITFILE cannot access a FITConnect save for sending in data queries and performing security updates. Data Controllers have configuration-level control over data requests, storage and releases

Key Node design features:

Deploy anywhere – any cloud any infrastructure.

Centralised control and deployment for easy authentication and update.

Containerised – each Node has every feature it needs to act independently, and self-heal when an issue is found.

A Node can act in a secure network of Nodes when controlled exchange of privacy treated private data is required.

Nodes are cost efficient and can be put to sleep when not in use.

Data Access

Each FITFILE Node contains runtime components and data stores. All data is encrypted at rest and transmitted with TLS encryption. 3 key components to facilitate the processing of local or remote data are FITConnect, HealthFILE and InsightFILE. Together, they perform access, privacy treatment, other processing and linkage.

is a modern software application deployed in the cloud or on premise inside the regulatory perimeter of the organisation data sources with required data elements such as EHRs, diagnostic and prescription databases

is for permitted users with the right legal basis or legitimate patient care interest to display identifiable or pseudonymised (linked) record-level data or insights

sends queries into FITConnect locally or remotely to identify populations stratified across data sources and display anonymised (linked) record-level data or insights

Data Privacy

FITFILE uses an automated protocol to privacy treat, integrate and work with harmonised data. 

A combination of techniques is configured to generate the agreed output and to ensure it is safely accessed and used.

Data x acquired from
Data Source A
  • Data Classification
  • Data Categorisation
  • Data Profiling
Privacy treatments and utility loss measurement
Re-id risk assessmentFail
Data is privacy treatedPass
Data y acquired from
Data Source B
  • Data Classification
  • Data Categorisation
  • Data Profiling
Privacy treatments and utility loss measurement
Re-id risk assessmentFail
Data is privacy treatedPass
Data integration
Privacy treatments and utility loss measurement
Re-id risk assessmentFail
Integrated data is
privacy treatedPass
Computation at source (using anonymised ciphers)
  1. Cohort of interest identified in Data Source A
  2. Irreversibly encrypted ciphers sent to Data Source B
  3. Overlap of records confirmed in Data Source B
  4. Overlapping ciphers sent back to Data Source A
  5. Any statistical attributes of interest calculated within each Data Source on the confirmed (overlapping and/ or non-overlapping) cohorts of interest
Harmonisation steps
  1. Identify the required data elements/documents
  2. Agree data quality dimensions and thresholds
  3. Define terms/ identify existing data dictionaries
  4. Select standard/ structured data model to
    • Consolidate data requirements
    • Align semantic definitions
    • Agree representation format

Data Linkage

Industry-standard linkage of pseudonymised data 1

Personal identifier(s) represented by unchanging artificial identifier

Useful privacy treatment

For research, planning or action purposes

Full linkability

More data can be easily added in future

1  Pseudonymisation or tokenisation is a privacy treatment procedure by which personally identifiable information is replaced by one or more artificial identifiers, typically using token-based approaches; additional info can restore the original (identifiable) state and hence, under the GDPR and common law, legal basis and/or approval are required

Globally unique linkage of anonymised data 2

Personal identifier(s) fully removed every time the process runs

Highest privacy

Individuals can never be re-identified

Highest efficiency

Faster, more cost-effective and reusable

2  Under anonymisation there is no key/ token to reverse the process, and (unlike pseudonymisation) individuals can never be re-identified

Anonymisation at source

ahead of linkage across sources

Why choose irreversible anonymisation over pseudonymisation?

Computation at source

for cohorts of interest

Why leave the data at source?

Want to go deeper?