VapusData OS
Governance belongs in the pipeline, not beside it
Catalogues describe data after it moves. Policy that travels with the pipeline is the only kind that can actually stop anything.
Vikrant Singh · · 3 min read
Most governance programmes start with an inventory. Someone stands up a catalogue, crawls the warehouse, and produces a searchable list of every table in the estate. It is genuinely useful, and it is also entirely descriptive: the catalogue learns that a column holds personal data some time after a pipeline has already copied it into three downstream systems.
The gap is not tooling. It is placement. A control that sits beside the pipeline can only ever report; a control that sits in the pipeline can refuse.
What "in the pipeline" actually means
Three properties, and a system either has them or it does not.
Policy is evaluated at the point of movement. Not on a nightly scan, not in a quarterly access review — at the moment a job asks for the data. If the answer arrives after the bytes have moved, it was an audit, not a control.
The decision travels with the data. A classification that lives only in the catalogue stops at the catalogue's boundary. One that is attached to the dataset survives the copy into a feature store, a notebook, or a model's training set.
Refusal is a normal outcome. This is the uncomfortable one. A governance layer that has never blocked a job is not proof that everything is compliant; it is proof that the layer is advisory.
The shape of the check
The useful mental model is a policy decision that returns before the read, not a log line written after it:
# Every read is brokered. The policy engine sees the purpose, not just the principal.
result = catalog.read(
dataset="finance.invoices",
purpose="ap-fraud-detection",
principal=agent.identity,
)The purpose field is doing more work than it looks. Access control that only knows who is
asking can express "the analytics team may read invoices". It cannot express "the analytics team
may read invoices for reconciliation, but not to train an external model" — and that second
sentence is what a data protection officer actually wants to enforce.
Where the residency question lands
Purpose-bound access and data residency turn out to be the same mechanism seen from two angles. If the broker knows which jurisdiction a dataset is bound to, it can decline a read from compute running elsewhere, rather than discovering the transfer in an egress log a month later.
| Approach | Detects a violation | Prevents a violation |
|---|---|---|
| Catalogue and scan | After the fact | No |
| Access review | On a cycle | Only for the next request |
| In-pipeline policy | At request time | Yes |
None of these replaces the others. A catalogue is still how a human finds a dataset, and access reviews are still how entitlements get pruned. But only the third row changes what is possible, and it is the row most estates are missing.
The trade you are making
In-pipeline policy costs latency on every read, and it will break jobs that were quietly relying on access nobody had approved. Both of those are real, and the second one is usually what stalls a rollout — the first month is spent discovering how much of the estate was running on informal access.
That discovery is the value. The alternative is not a cheaper system; it is the same discovery made later, by a regulator.
Read more about how this works in VapusData OS, or see the deployment topologies that keep the policy plane inside your own network boundary.
Tags
governance · architecture · data platform

