Integrating legacy databases with new processes via controlled data contracts
The new API returns a successful response, but the legacy database interprets a status, identifier, or date differently from the new process. The message is technically delivered, yet the business operation fails or silently distorts financial reporting. Resolving this requires more than network connectivity: the systems must also be semantically interoperable.
Why technical connection does not mean data compatibility
In system integration practice, there is often an illusion that successfully establishing a network connection or agreeing on transmission formats (for example, JSON via REST) automatically solves the interaction problem. However, technically connecting systems at the network or protocol level does not mean semantic compatibility of their data. Each system has its own domain model, which has been shaped over the years under the influence of specific business requirements and the limitations of past technological generations.
When a new service attempts to interact with a legacy database directly, a conflict of interpretations arises. For example, a status field in the new system may have strict typing and transition logic, whereas in the old database it is stored as an arbitrary text string or numeric code without documentation. Attempting to map such fields directly leads to an accumulation of errors that are difficult to detect at the infrastructure monitoring level.
Scope and limits of OpenAPI
To describe the interaction between modern services, using the OpenAPI specification has become standard practice. It defines a programming language-agnostic description of the HTTP API interface, enabling humans and computers to understand the capabilities of the service without analysing its source code. This is an effective tool for documenting a technical contract, but an architect must consider its limits.
Describing an interface with OpenAPI does not guarantee data security or consistency. A Schema Object can define structure, types, and local value constraints such as required, enum, ranges, or pattern, but enforcement depends on validator and runtime tooling. Domain logic must check rules that depend on current system state or interactions between systems. A schema can require a non-negative number, but it does not know the actual stock balance. OpenAPI can also describe security schemes and requirements, but the document itself does not implement authentication, authorisation, or transactional consistency.
Isolating legacy models through an Anti-Corruption Layer
To prevent legacy semantics from leaking into new business processes, architects use the Anti-Corruption Layer (ACL) pattern. This layer isolates different system models and semantics through a dedicated facade or adapter. Instead of adapting a new microservice to the limitations of an old database, the ACL translates requests between the canonical and legacy formats.
Designing such an adapter requires trade-offs. An Anti-Corruption Layer adds translation logic and latency; it can be implemented as an in-application component or an independent service, and interaction through it can be synchronous or asynchronous. Regardless of deployment shape, the ACL needs dedicated testing, monitoring, scaling, and maintenance. At the integration boundary, validate and sanitise input, control schema evolution and consistency, use correlation IDs and structured logs, and plan for failure of the adapter itself.
Roles of technical approaches in an integration pipeline
These approaches operate at different architectural layers and can be combined: an ACL isolates domain semantics, logical replication and CDC capture and transmit changes, and event-driven exchange defines asynchronous interaction. Choose the combination according to latency, volume, consistency, recovery, and data-loss tolerance.
| Approach | Mechanism | Main limitations and risks |
|---|---|---|
| Anti-Corruption Layer | Synchronous requests or asynchronous messages through a semantic adapter | Latency, translation errors, scaling, availability, and maintenance |
| PostgreSQL logical replication | Publication/subscription: initial copy followed by a stream of changes |
Conflicts or overwrites from local writes; DDL, sequence state, and large objects are not replicated |
| Change Data Capture, such as Debezium | Capturing changes from a log or replication stream and forwarding them through Connect, Server, or Engine | Replication lag, duplicates under at-least-once delivery, schema evolution, source-log retention, and monitoring |
| Event-driven exchange | Asynchronous messages through a broker or transport; CloudEvents can provide the event envelope | Delivery, duplication, ordering, and exactly-once behaviour depend on the entire path and its configuration |
In a typical PostgreSQL logical-replication configuration, table state is copied first, after which changes are continuously transmitted in commit order within a single subscription. An incoming UPDATE can overwrite a locally modified row; an UPDATE or DELETE whose target row is missing is skipped; a constraint violation stops the apply worker until the conflict is resolved. Schema and DDL, sequence state, and large objects are not transferred by logical replication and must be synchronised separately.
Debezium builds a CDC pipeline through connectors to database change logs and can run through Kafka Connect, Debezium Server, or the embedded Debezium Engine. Its baseline guarantee is at-least-once: after a failure or restart, the same change can be delivered again. Consumers therefore need to be idempotent or deduplicate events by source position, LSN, or another stable identifier.
CloudEvents is often used to standardise event descriptions. It has protocol bindings including HTTP, Kafka, AMQP, and MQTT, and separate event formats including JSON and Avro. It defines how an event is represented, but it does not establish end-to-end transfer or settlement rules. Delivery, duplication, and ordering guarantees depend on the producer, partitioning, transport or broker, consumer configuration, offset commits, and the target write. Handlers should therefore generally be idempotent or perform deduplication.
Strategies for conflict control during parallel writes
When systems operate in parallel, first define the write topology: whether both paths update one authoritative store or maintain independent writable stores, which system is the source of truth, and whether a consumer applies a patch, full-row upsert, or merge. Typical CDC pipelines transmit changes asynchronously and therefore introduce replication lag, but lag alone does not define how a conflict is resolved.
Optimistic locking works when a command carries a version or revision token read with the data and the authoritative writer atomically compares it with the current version; every write path must follow the same rule. Two independent writable stores additionally require single-writer ownership or an explicit protocol for detecting, merging, and resolving conflicts. Use a wall-clock timestamp as a token only when its semantics, precision, and time source are guaranteed.
For critical data, retain audit and provenance metadata for each change: the source system, operation or event ID, source position or version, and change time. Choose record-level or field-level granularity according to risk and regulatory requirements. Data lineage separately describes the origin and transformations of data across datasets, jobs, and fields.
Building integration layers based on UnityBase
IQusion offers custom solution development and integration with existing information systems and uses the UnityBase platform for enterprise web-oriented systems. Its official documentation describes an HTTP(S) server, a DBMS-agnostic ORM over domain metadata, and REST API generation. These capabilities allow a team to implement integration adapters and ACLs for a specific solution architecture, but they are not a ready-made ACL framework and do not by themselves eliminate data-integrity risks.
- Semantic mapping: For every field, document transformation direction, precision, and possible information loss. If reverse conversion is ambiguous, retain the original value and provenance and route unknown values to quarantine or manual review.
- Model isolation (ACL): Business commands and updates from the new domain should cross the controlled boundary. Document read replicas, CDC readers, reporting, or migration tooling as special paths and do not let them become uncontrolled alternative write paths.
- Write ownership and conflicts: Define the source of truth and write owner for each entity or field. Specify version-token checks where applicable, together with merge, rejection, and manual-resolution rules.
- Delivery and idempotency: Specify end-to-end delivery semantics, partitioning keys, redelivery and deduplication rules, and verify that replaying an event does not create another business effect.
- Lag objective: Define an end-to-end replication-lag SLI and SLO, a logical watermark for measurement, and an alert threshold; use SLA only when a contractual guarantee exists.
- Reconciliation: Compare systems at a common logical cutoff or watermark. Use record count only as a coarse indicator; compare keys, versions, and normalised row-level or range-level checksums and report missing, extra, and different records separately.
Frequently Asked Questions
Can OpenAPI be used to guarantee data security during integration?
No. OpenAPI describes an HTTP API interface, data schemas, some local value constraints, and security schemes and requirements, but enforcement depends on runtime tooling. Authentication, authorisation, current business-state validation, and transactional consistency must be provided by the system implementation.
What are the main disadvantages of using an Anti-Corruption Layer (ACL)?
An ACL adds translation logic and latency, can become a point of failure, and requires dedicated testing, monitoring, scaling, and maintenance. Depending on the architecture, it can be an in-application component or an independent service and can support synchronous or asynchronous interaction.
Does using CloudEvents guarantee message delivery?
No. CloudEvents standardises an event's structure and representation, but it does not define end-to-end delivery, deduplication, ordering, or exactly-once processing guarantees. These depend on the producer, transport or broker, partitioning, consumer, and target write, so handlers should generally be idempotent.
Sources
- OpenAPI Specification 3.2.1
- Microsoft: Anti-Corruption Layer pattern
- PostgreSQL 18: Logical Replication
- PostgreSQL 18: Conflicts
- PostgreSQL 18: Restrictions
- Debezium 3.6 Architecture
- Debezium 3.6 Delivery Guarantees
- Debezium 3.6 PostgreSQL Connector
- CloudEvents 1.0.2 Document Index
- CloudEvents 1.0.2
- CloudEvents Kafka Protocol Binding
- Microsoft: Optimistic Concurrency
- W3C PROV Overview
- OpenLineage: Dataset Lineage Facet
- AWS DMS Data Validation
- UnityBase Documentation
- Google SRE: Service Level Objectives
- Software