Several AI agents on the left, a lake, warehouse, and vector store on the right, and a single layer between them labelled with policy, keys, recovery, lineage, and discovery

I describe Datris as the data control plane for AI agents, and a few people have asked me what that actually means. Fair question. We didn’t start out building a control plane. We shipped scoped keys, then an approval policy, then a recovery loop, then lineage and a catalog, and somewhere along the way I realized they were all parts of the same layer.

Where the phrase comes from

Network engineers split a router in two a long time ago. The data plane moves packets at line rate. The control plane holds the routing table, the policy, and the identity of every peer, and decides where packets go without touching one. Kubernetes borrowed the split, and its control plane is the reason a hundred people can deploy to the same cluster without a meeting. Nobody argues that a cluster doesn’t need one. The argument I keep having is whether data does, now that the operator is an agent.

The operator changed and the stack didn’t

For thirty years the data stack assumed the operator was a person who’d read the runbook, knew which table you don’t touch on month-end, and remembered what they broke. Every safety property we have leans on that. An agent has no memory of yesterday, month-end means nothing to it unless something says so, and it runs at 3am on a Sunday with the same confidence it has at noon on a Tuesday. It’s also fast, so a wrong turn becomes eight million wrong rows before anyone’s coffee is ready.

I’m not interested in locking agents out. I’ve watched one stand up a pipeline, provision its own credentials, build a tap, run it, and query the result, and once you’ve seen that you don’t go back to the ticket queue. The properties we used to get from the person have to come from somewhere else, and that somewhere is the control plane.

What it has to answer

A data control plane sits between every agent and every store and answers a handful of questions in the request path, where the agent has no say.

What may this agent do at all? The credential carries the answer. One broad API key per deployment means every agent is every agent, so the control plane issues scoped keys that bind a capability list to the key.

What may it do alone? A key is yes or no, and a delete or a schema migration needs “yes, but a person looks first.” So an administrator sets a policy the agent can’t read, marking each action auto, approve, or deny. Approved requests are captured, parked on a card, and run as proposed only when a human clicks. The agent can never edit the policy or approve its own request.

What does it do when things break? Most demos stop at the first successful tool call, which is roughly where production gets interesting. The control plane watches runs, and when one fails it proposes the repair through the same policy. You earn autopilot one action at a time.

What did it touch? You need which run, which agent, acting for whom, under which approval. That means an audit log and column-level lineage populated automatically, because the agent won’t recall what it did and you can’t depose it.

What exists? An agent that only sees its own pipelines is lost in ten years of landed data. The control plane holds one catalog across structured, semi-structured, and unstructured data, so the agent finds a dataset before it rebuilds it.

Underneath all of these sits the question of secrets. The agent works the lock and the vault holds the key, so credentials rotate without any agent noticing and no agent can leak what it never held.

Every bank I worked in had all of this, for people. The difference now is that it has to be mechanical.

Why it isn’t the warehouse or the framework

The vendors would like this layer inside the lake, and they’re building toward it. The trouble is that the agent’s job starts earlier, at a source nobody wrote a connector for, with a key fetched from somewhere, under a policy that applies before any row lands. A control plane inside the lake only governs the second half of the story. Ours sits beside the lake and treats the lake, the warehouse, and the vector store all as destinations.

Putting it in the agent framework has a different problem: you don’t control which framework shows up. This year it’s a desktop client, an IDE plugin, a midnight shell script, and three orchestrators from three vendors. A control plane that only covers the agents you wrote is a prompt with extra steps. That’s why ours speaks MCP, the one door every agent already knows how to knock on.

The upshot

What surprised me is that a good control plane makes the agent simpler. Once it knows it’ll be refused, parked, or told exactly which column it broke, it can stop narrating defensively and get on with the work. The caution lives in the layer, which is good, because prompts are where caution goes to be forgotten.

Datris is one of these. It’s open source and runs beside whatever lake you already have. The argument doesn’t depend on my product, though. If your agents can reach your data and nothing in between holds the rules, you’ve got a cluster with no API server, and it’ll be fine until the first bad Sunday.


Todd Fearn is the founder and CEO of Datris, the open-source data control plane for AI agents, built on the Model Context Protocol. He also runs IData Corporation, the data engineering consultancy Datris grew out of. Before that he spent thirty years building production data infrastructure at Goldman Sachs, Bridgewater, Deutsche Bank, Salomon Brothers, Freddie Mac, and others, and founded several venture-backed startups.