Data Extraction from Silos – the CODE24 Approach

Healthcare IT is at a crucial crossroads. Everyone is talking about data availability, regional collaboration and data-driven care. However, the day-to-day reality for IT departments is often considerably more challenging. Data is frequently stuck in silos, locked away in healthcare systems that work perfectly well for a specific department or organisation, but are largely inaccessible to the outside world. At CODE24, we believe the true value of healthcare data lies in its ability to flow freely and be reused. Why should data have to be manually entered again and again, or copied countless times to make it available in other systems?

To become truly future-proof and achieve genuine data availability, data needs to be extracted from traditional systems and converted into a future-proof, vendor-neutral format: openEHR. By modelling and storing data according to a generic, open information model, you create a solid foundation on which countless new applications can be built. But how do you do this when healthcare IT vendors do not support – or even actively obstruct – this approach?

At CODE24, we have extensive experience in unlocking data from a wide range of healthcare systems. We are happy to share what we have learned.data-extractie-uit-silo's-door-CODE24

The challenge: the door is closed

If you have been working in healthcare IT for a while, you will recognise the pattern immediately. Legacy healthcare systems have a number of persistent characteristics that make data extraction challenging. They often rely on a strictly vendor-dependent information model, while database schemas are frequently unavailable.

Building connections between systems to enable interoperability can also be tricky. For every new integration, use case or desired functionality, new API connections often need to be developed – costing healthcare organisations significant amounts of time and money.

There are therefore several barriers to overcome – so let’s break them down.

The approach: structured extraction

How do you approach a data extraction project like this without getting stuck in a multi-year plan? The key lies in robust governance combined with a pragmatic extraction plan. You need to know exactly what you are going to extract, where it is located and how often you need to do so. In most cases, we are talking about one-way data flows: we read data from the legacy healthcare system and write it to openEHR. Data is generally not written back to the legacy source, unless the vendor provides the necessary means to do so, such as an available API.

The six-step information extraction plan

Before the technology starts doing its work, we define a clear plan based on six crucial questions:
  1. Which healthcare system? Identify the exact legacy source, including the specific software version, from which the data originates.
  2. Which dataset? Define the scope. There is no need to blindly copy everything. Which clinical data is actually relevant to the use case?
  3. Which method? Determine the technical approach based on the options available (see the overview of methods below).
  4. How frequently? How up to date does the data in openEHR need to be? Does it need to be available immediately (near real-time), or is every minute, hour, day or week sufficient?
  5. Which information object? Clearly define the metadata, exact storage location and version control of the target element.
  6. Which process? Establish ownership, error handling and monitoring around the data flow (who steps in when an extraction fails)?

the method

When it comes to the extraction method, we always look at what is available and which option will deliver the best results. We use the following order of priority:

  • The API: If a usable, documented API is available for the healthcare system, we naturally use this as the formal route.
  • Extraction via messages: If there are no direct APIs, but the system sends out FHIR resources, HL7 v2 messages or CDA documents in response to certain events, we capture these data flows to feed the openEHR environment.
  • Direct SQL queries: If the system is completely closed and there is no support available from the vendor, it may still be possible to extract the required tables directly from the underlying database using SQL queries.
    The data warehouse: If a data warehouse containing standardised data from the healthcare system is already available, we can extract the data directly from there.

The role of the ETL tool

To manage this well-defined process, we use a robust ETL (Extract, Transform, Load) tool. The tool needs to be able to work seamlessly with the different (FHIR) dialects used by vendors, because unfortunately, ‘standard FHIR’ is rarely truly standard in practice. The ETL tool must also be able to run reliably and unobtrusively as a background process, with the extraction frequency configurable independently for each connected healthcare system.

Lessons learned from practice

All this theory sounds great, but how does it translate into the realities of day-to-day healthcare IT? At CODE24, we have successfully applied this approach to a range of complex use cases, including:

  1. Making data available from HiX to a PGO
    Making structured data from an EHR such as ChipSoft HiX available to a MedMij-certified Personal Health Environment (PGO) can often be a major challenge. By synchronising data using HL7 and retrieving additional data through queries, we were able to make the data available through an open healthcare data platform and subsequently transfer it via the PGO’s API. A great example of how data can be made available directly to the patient.

  2. Shared treatment for patients with severe mental health conditions
    When treating patients with severe mental health conditions, multiple mental healthcare organisations may work together. Two different organisations, different EHRs, but one shared patient. By securely bringing the data together in an openEHR data platform through our extraction platform, clinicians at both organisations gain access to a single, up-to-date shared view of the patient’s treatment.

  3. Outpatient collaboration with active data access
    In outpatient care, you do not want to find out what happened yesterday – you need the data now. For one of our customers, we implemented active data access, continuously extracting the most up-to-date data from the source EHR for use in an outpatient application. Here too, we used a combination of HL7-based synchronisation and retrieving healthcare data through queries.

Healthcare data autonomy starts today

The ability to extract data from healthcare systems and standardise it for multiple uses is a strategic necessity for healthcare organisations that want to regain control over their own healthcare data. It is not easy. But with the right extraction plan, a flexible ETL layer that understands different dialects, and a clear vision for data reuse, we have repeatedly demonstrated that it can be done.

Curious about how we can get your healthcare data out of its silos and ready for multiple uses? As a Value Added Reseller (VAR) of the openEHR data platform Cadasto, we have both the expertise and the platform to help you successfully get such a project off the ground.

Digitalising Your Lab Ordering Process: 5 Lessons Learned from Real-World Practice

meet: Product Owner Lieke