Nobody is being dishonest. That is what makes it a trap rather than a lie.
"We encrypt it and sign a data-processing agreement" is a real control. It is
not the same kind of thing as processing that never leaves the building, and the
difference only becomes visible at the moment it matters.
The phrase that hides the work is permitted subject to controls. Almost
nothing in this space is flatly forbidden. Things are allowed, provided you take
on a set of obligations — and the set is presented as a checklist, which makes
it look like a project with an end.
It is not a project. It is a subscription. Here is what you are actually
signing up for.
This post describes how controls behave, not what any regulation requires.
Your counsel owns the second question and it deserves its own conversation.
The stack, laid out
To run a third-party model over sensitive material, an organisation typically
executes and maintains some version of the following: a data-processing
agreement with the provider and, transitively, its sub-processors; an assessment
covering where processing occurs; a documented purpose for each use; encryption
in transit and at rest with managed key custody; a contractual assurance that
the provider will not train on the data; notification chains that meet whatever
clock applies to you; retention and deletion controls that can be demonstrated;
and audit rights exercisable only through the provider's own reporting.
Every item on that list is achievable. Most organisations can produce all of
them. The trap is not that any single control is hard. It is three properties
they share.
Property one: they recur
Each control must be re-performed whenever the provider changes a
sub-processor, adds a region, deprecates a model or updates its terms.
None of those events is under your control, and all of them are routine for a
platform serving millions of customers. A sub-processor is added — your
disclosure is now incomplete. A model is deprecated — your validation covered
the old one. Terms are updated — your assessment referenced the previous
version.
The control is not a thing you did. It is a review cycle you now own, running at
a cadence somebody else sets.
The practical failure is not dramatic. It is that the second cycle is done
thoroughly, the fourth is done quickly, and the seventh is done by someone who
inherited the folder and does not know why any of it is there.
Property two: they are exercised at arm's length
Your audit right is a right to receive a report about a system you cannot
inspect.
This is a real control and it should not be dismissed. It is also categorically
different from looking. What you are receiving is an assertion produced by the
party being audited, in a format they chose, at a frequency they set.
Ask what you would do if a report were incomplete, and the answer is: request
clarification. That is the whole enforcement mechanism, and it is why the
distinction between a control you hold and a control you are told about is worth
keeping in view.
Property three: they produce assurance, not evidence
At the end of the process you hold a contract stating that a thing did not
happen. What a reviewer asks for is a record of what did.
This is the property that matters most, because it is invisible until you need
it. A stack of well-maintained compensating controls answers the question was
this permitted? It does not answer what happened to this record on the
fourteenth of March?
For the second question you need an artefact produced at the time, by the
system, as a side effect of doing the work. Either the architecture emits one or
it does not, and no amount of contractual diligence retrofits it.
What "removing the question" actually means
The alternative is not a cheaper way of satisfying those controls. It is an
architecture in which most of them have nothing to attach to.
There is no transfer, so there is no transfer to assess. There is no
sub-processor, so there is no chain to disclose and re-disclose each quarter.
There is no third-party training question, because the weights and the data sit
on hardware you own.
This is a smaller claim than "on-premise makes you compliant," and it is a
true one. It does not eliminate obligations. It eliminates a category of
obligation that exists only because a third party is in the processing path,
and leaves you with the rest: access control, purpose limitation, minimisation,
retention, auditability. You were already doing those for every other system in
your estate, and you already know how.
Where this argument does not apply
Being cautious about compensating controls is not a reason to avoid cloud
services generally, and we are not going to pretend otherwise.
If the material is not sensitive, the control stack is thin and the trap does
not spring. If the volume is low, the arithmetic of owning infrastructure does
not work regardless of the governance argument — below roughly 90,000 queries a
month, no on-premise deployment we build pays back inside three years, including
our entry tier.
And if the task genuinely needs frontier-level capability on non-sensitive
material, running something smaller locally is the wrong instrument whatever the
control story.
What we most often recommend is a split — sensitive and high-volume work on
owned infrastructure, everything else on cloud services, with the routing
enforced in code rather than by staff discipline. The reason to enforce it in
code is not distrust of employees. It is that a policy-based split produces no
evidence that it held, which puts you back at the top of this post.
The one question that exposes the trap
"If I am asked what happened to this specific record eight months ago, what
do I produce?"
If the answer is a contract, you have assurance. If the answer is a record, and
you can show it has not been altered since, you have evidence.
Both are positions. Only one of them survives the question.