Step-by-Step Modernisation of Critical IT Systems: Risk Management, Metrics, and Rollback Scenarios

October 2, 2026 · 8 min

The new component is already prepared to receive traffic, but the data schema, dependencies, and rollback procedure to the old version have not yet been verified together. Architects of critical systems often find themselves at this point of uncertainty, where business pressure demands rapid changes, while the risk of failure threatens operational activities. Modernisation of large infrastructure solutions rarely forgives a sudden replacement approach, which is why an evolutionary transition becomes the only manageable path to preserve the stability of business processes.

Defining the boundaries of the first stage and isolating functionality

Successful modernisation begins with decomposition. Attempting to replace the entire system at once creates too many unknown variables, making the search for the causes of potential failures virtually impossible. The primary task of the architect is to isolate the least coupled, yet representative function of the legacy system for initial replacement. This function must have clearly defined data boundaries and a minimal number of synchronous dependencies on other modules.

When choosing the first candidate for replacement, one should not touch the critical core of the system, but at the same time, one should not select overly trivial services that will not demonstrate real integration problems. The optimal choice is an isolated business process with a moderate load. Isolating this function allows the team to work through the entire lifecycle of deployment, monitoring, and interaction between the old and new parts of the system with minimal risk to the organisation's core activities.

The Strangler Fig pattern and router risk management

For the gradual replacement of legacy system functionality, the Strangler Fig pattern is used. Its essence lies in the gradual redirection of requests from old components to new ones through a special interface layer — a facade. The facade intercepts incoming calls and routes them depending on the readiness of the new implementation. Over time, the new service completely replaces the old functionality, and the legacy component is decommissioned.

However, the facade itself becomes a new critical point of risk and a potential bottleneck of the entire architecture. If the router fails, the entire system stops. Therefore, designing the facade requires special attention to its fault tolerance, scalability, and simplicity of logic. It must not contain complex business logic — its sole task is to redirect traffic quickly and reliably.

To minimise such architectural risks, IQusion uses the UnityBase low-code platform as a technological foundation. Thanks to the model-driven approach, where the domain model automatically generates REST/ORM-APIs and user interfaces, the volume of manual code is significantly reduced. This simplifies the creation of reliable facades for the Strangler Fig pattern, the configuration of integrations via open APIs, and the updating of individual modules without stopping the entire system, while ensuring strict access control and a detailed audit of actions.

Stability metrics: distinguishing between SLI, SLO, RTO, and RPO

A safe transition is impossible without clear quantitative criteria for evaluating system performance. Architects must clearly distinguish between service level indicators (SLI) and service level objectives (SLO). An SLI is a specific quantitative measurement of service behaviour, such as response time or the proportion of successful requests. An SLO is a target value or range for this indicator, agreed with the business. Absolute one hundred per cent availability is not a realistic goal, as the costs of achieving it grow exponentially without bringing proportional value.

In addition to operational metrics, business continuity indicators are critically important: recovery time objective (RTO) and recovery point objective (RPO). RTO defines how long a system can be unavailable during an incident before its full recovery. RPO indicates the acceptable volume of data loss, measured in time from the moment of the last saved state.

The success of modernisation is measured not by the speed of writing code, but by the system's ability to return to a stable state within the defined RTO and RPO when any anomaly occurs.

Why state migration requires a separate scenario

Popular safe deployment strategies, such as Canary or blue-green, work excellently for stateless components, but they do not automatically solve the complex problems of schema and data state migration. If a new version of a service requires changes to the database structure, a simple traffic switch can lead to incompatibility and data loss during a rollback attempt.

For stateful components, a separate scenario for synchronisation and backward compatibility support is required. A typical approach is to use a dual-write strategy, where data is simultaneously written to both the old and the new databases. This approach keeps both repositories in a consistent state and ensures the ability to roll back to the old version without losing transactions completed during the testing of the new component.

Quality gates and launch readiness assessment algorithm

Before expanding the volume of traffic to a new component, it must pass through a system of quality gates. These gates determine whether the current state of the system meets the established safety and stability criteria.

  • Test coverage assessment: Verifying the presence of integration and regression tests for the new component (minimum 80% coverage of critical paths).
  • Performance metrics validation: Comparing the SLI of the new component on the test environment with the current SLOs of the legacy system under target load.
  • Data rollback plan verification: Conducting a failure simulation with data state recovery to a defined RPO point.
  • Facade error monitoring: Verifying the absence of an abnormal increase in 5xx errors at the router level when redirecting 1% of traffic.
  • Support service readiness: Availability of instructions for operators and configured incident notification dashboards.

Controlled rollback scenario upon anomaly detection

If a deviation of SLI from target SLO metrics is detected during gradual traffic switching, a clear rollback scenario must be activated. This process must not be a chaotic decision by an on-duty engineer — it is a pre-automated and tested algorithm of actions.

The first step is to instantly return routing on the facade to 100% usage of the old system. After isolating the new component, a data state analysis is conducted: if transactions were recorded that managed to write only to the new database, a data reconciliation and compensation procedure is launched to restore integrity in accordance with the established RPO. Only after the full restoration of the stable state of the legacy system does the team proceed to log analysis and troubleshooting the causes of the incident.

Frequently Asked Questions

What is the Strangler Fig pattern and how does it help during modernisation?

The Strangler Fig pattern involves the gradual replacement of individual functions of a legacy system with new components. A special interface layer (facade) intercepts requests and routes them between the old and new implementations, reducing the risk of a large-scale failure during the transition.

Why does a Canary deployment not solve the data migration problem automatically?

A Canary deployment only manages the distribution of user traffic, but not the structure of stored data. Changing data schemas in stateful components requires separate synchronisation and backward compatibility scenarios to ensure a safe rollback capability.

What is the difference between RTO and RPO metrics?

RTO (Recovery Time Objective) defines the acceptable duration of system recovery after a failure. RPO (Recovery Point Objective) indicates the acceptable point in the past to which data must be recovered without critical losses for the business.

Sources