VapusData logoVapusData logo
← All posts

VapusData OS

Governance belongs in the pipeline, not beside it

Catalogues describe data after it moves. Policy that travels with the pipeline is the only kind that can actually stop anything.

Vikrant Singh · · 3 min read

Most governance programmes start with an inventory. Someone stands up a catalogue, crawls the warehouse, and produces a searchable list of every table in the estate. It is genuinely useful, and it is also entirely descriptive: the catalogue learns that a column holds personal data some time after a pipeline has already copied it into three downstream systems.

The gap is not tooling. It is placement. A control that sits beside the pipeline can only ever report; a control that sits in the pipeline can refuse.

What "in the pipeline" actually means

Three properties, and a system either has them or it does not.

Policy is evaluated at the point of movement. Not on a nightly scan, not in a quarterly access review — at the moment a job asks for the data. If the answer arrives after the bytes have moved, it was an audit, not a control.

The decision travels with the data. A classification that lives only in the catalogue stops at the catalogue's boundary. One that is attached to the dataset survives the copy into a feature store, a notebook, or a model's training set.

Refusal is a normal outcome. This is the uncomfortable one. A governance layer that has never blocked a job is not proof that everything is compliant; it is proof that the layer is advisory.

The shape of the check

The useful mental model is a policy decision that returns before the read, not a log line written after it:

# Every read is brokered. The policy engine sees the purpose, not just the principal.
result = catalog.read(
    dataset="finance.invoices",
    purpose="ap-fraud-detection",
    principal=agent.identity,
)

The purpose field is doing more work than it looks. Access control that only knows who is asking can express "the analytics team may read invoices". It cannot express "the analytics team may read invoices for reconciliation, but not to train an external model" — and that second sentence is what a data protection officer actually wants to enforce.

Where the residency question lands

Purpose-bound access and data residency turn out to be the same mechanism seen from two angles. If the broker knows which jurisdiction a dataset is bound to, it can decline a read from compute running elsewhere, rather than discovering the transfer in an egress log a month later.

ApproachDetects a violationPrevents a violation
Catalogue and scanAfter the factNo
Access reviewOn a cycleOnly for the next request
In-pipeline policyAt request timeYes

None of these replaces the others. A catalogue is still how a human finds a dataset, and access reviews are still how entitlements get pruned. But only the third row changes what is possible, and it is the row most estates are missing.

The trade you are making

In-pipeline policy costs latency on every read, and it will break jobs that were quietly relying on access nobody had approved. Both of those are real, and the second one is usually what stalls a rollout — the first month is spent discovering how much of the estate was running on informal access.

That discovery is the value. The alternative is not a cheaper system; it is the same discovery made later, by a regulator.

Read more about how this works in VapusData OS, or see the deployment topologies that keep the policy plane inside your own network boundary.

Tags

governance · architecture · data platform

Essential cookies are required for the site to function and cannot be switched off. Everything else is off until you switch it on, and you can change or withdraw your choice at any time from the Cookie settings link in the footer. The Cookie Policy lists the cookies we set and how long each one lasts.

No choice recorded yet