GDPR Erasure vs. Immutable Event Logs
Reconciling Right-to-Be-Forgotten with Event Sourcing
Reconciling Right-to-Be-Forgotten with Event Sourcing
"An append-only event log and the right to be forgotten are fundamentally incompatible" is the sentence that ends up in half the architecture reviews I've sat in for an event-sourced ledger, and it's wrong in a specific, useful way: it treats every field in every event as equally subject to erasure, when for a bank, most of what's actually in a transaction event isn't erasable in the first place — not because of an architectural workaround, but because GDPR itself says so. Scoping that correctly is the actual first step, and it's a BA question before it's an engineering one.
The exception that changes the whole spec
Article 17(3)(b) of the GDPR exempts a controller from erasure where retention is necessary to comply with a legal obligation — and for a bank, that's not a theoretical carve-out. AML record-keeping rules, MiFID II transaction reporting, and national commercial-record retention laws (Germany's HGB requires ten years for accounting records; most EU jurisdictions sit in the six-to-ten-year range) mean the core facts of a transaction — who transferred what, to whom, when, for how much — are legally required to be retained regardless of an erasure request, for as long as the applicable retention period runs. A customer can ask, and the bank's honest, defensible answer for that data is "no, and here's the regulation that says why," communicated in writing within one month per Article 12(4), not silence, and not a workaround engineered to technically comply while functionally not.
That reframes the actual acceptance criterion, and it's worth writing explicitly rather than leaving implicit:
While a transaction event falls within the applicable AML or financial
record retention period, the system SHALL NOT erase or irreversibly
obscure the core transaction facts (parties, amount, timestamp,
instrument) required to satisfy that legal obligation, and SHALL
respond to any erasure request covering that data by citing the
specific retention obligation relied upon.
When an erasure request concerns personal data within an event that
is NOT itself required for the legal retention obligation — a free-
text memo field, a session IP address, a marketing consent flag
captured alongside the transaction — the system SHALL erase or
irreversibly anonymize that specific data, independent of the
retention status of the transaction event it's attached to.
That second clause is where the actual engineering problem lives, and it's narrower than "make the whole event log erasable" — which is exactly why scoping it correctly first saves a team from over-building a mechanism for data that was never erasable to begin with, or worse, under-building one and leaving genuinely erasable PII stuck inside an immutable structure by accident.
Pattern one: PII externalization
The cleaner pattern, where it fits: don't put personal data in the event payload at all. Reference the data subject by an opaque, stable identifier, and store the actual PII — name, address, contact details, free-text notes — in a separate subject-record store keyed by that identifier. The event log stays exactly what event sourcing wants it to be: an immutable, replayable sequence of facts, referencing a subject ID rather than containing the subject's actual details. Erasure becomes a normal, mutable-table delete or anonymization against the subject-record store — no rewriting of history required, because the history never held the PII in the first place.
The caveat worth flagging explicitly to whoever's reviewing this against legal sign-off: externalizing PII into a separate, referenceable table is pseudonymization, not erasure, for as long as that table and the mapping between subject ID and identity still exist — pseudonymized data is still personal data under GDPR. The erasure event has to be the actual deletion of the subject record, not the architectural decoupling itself; a design that stops at "we moved the PII to a different table" without a deletion mechanism for that table hasn't implemented erasure, it's implemented a place where erasure could later happen.
Pattern two: crypto-shredding
Externalization works cleanly when the PII is genuinely separable metadata — a contact address doesn't change what a transfer event means when replayed. It doesn't work as cleanly when the personal data is structurally part of what the event actually represents: a KYC decision event that recorded the customer's name and address as evaluated at the time, where replaying the decision logic later meaningfully depends on what those values were, not just a reference to who they belonged to.
For that case, crypto-shredding is the pattern that holds up: encrypt the PII-bearing portion of the event payload with a key unique to that data subject, and on a valid erasure request, destroy the key rather than the event. The event stays physically and structurally intact — append-only, replayable, position in the log unchanged — but the personal data it once held decrypts to nothing. This is genuinely useful for a banking ledger's backup and archive tiers specifically, where selectively deleting one subject's data out of a petabyte-scale backup is the "manifestly disproportionate effort" scenario Article 17(1) itself contemplates as a legitimate reason not to attempt literal deletion.
It comes with a caveat that belongs in the spec, not left for engineering to discover during a compliance review: the EDPB has not formally endorsed crypto-shredding as satisfying Article 17 erasure on its own, though several national data protection authorities have accepted it specifically where the alternative — reconstructing and selectively purging immutable backups — is disproportionate. That means crypto-shredding needs the same legal sign-off before implementation that a retention-obligation refusal needs, documented as a deliberate design decision — the key-destruction event logged and evidenced — not assumed to be self-evidently compliant because it's technically elegant.
The decision, and who actually signs off on it
The choice between the two patterns isn't really an engineering preference; it's a property of the data. Externalize a field when it's independent of the event's business meaning and only ever needed as a reference — contact details, marketing preferences, free-text notes attached to but not load-bearing within the transaction. Crypto-shred a field when the personal data is structurally embedded in what the event represents and replay integrity genuinely depends on the value being present, even if inaccessible to a human reader after the key is gone.
Either way, the retention-obligation scoping has to happen first, and it's a legal and compliance decision as much as an architectural one — the BA's job is delivering a spec that draws the line between "legally required to retain, full stop" and "erasable, using one of two patterns depending on how the data is structurally used," before an engineer designs a mechanism for data that a straight reading of Article 17(3)(b) says was never going to be erased in the first place.