Two lanes running side by side: the warehouse lane for analytics at scale, and a neutral ingestion-and-quality lane that loads it — same pipeline able to feed Snowflake, Databricks, Postgres, MongoDB, or object storage without switching allegiance

Data has gravity. People in this industry repeat that like it’s a law of physics, and as far as it goes, it’s true: data piles up somewhere, moving it costs money, so the tools drift toward wherever it sits.

But notice who says it loudest. The warehouse vendors. They aren’t describing gravity. They’re selling it.

You know the pitch. You already store your data with us, so why not transform it here too? Run quality here, build your ingestion here, govern here, run your agents here. Every step is reasonable on its own, and every step saves you an integration this quarter. Meanwhile a little more of your leverage moves inside somebody else’s walls, until the tool that feeds your warehouse, the logic that cleans it, and the agents that read it all belong to the same company that stores it.

I spent thirty years building data infrastructure inside financial institutions, and I watched that cycle run more times than I care to remember. Lock-in never announces itself as lock-in. It shows up wearing convenience. You find out what it actually was at renewal time.

The intake valve

Here’s the tell. Take the tool that loads your warehouse and ask one question: does it work anywhere else?

If your ingestion only lands in one vendor’s tables, if your quality rules only run on one vendor’s compute, if your transformations are written in one vendor’s dialect, then you don’t have a pipeline. You have an intake valve owned by the destination, and its whole purpose is to make sure your data arrives in the one place it’s hardest to leave.

You also pay for it twice: once in the captivity, once in the compute. When quality runs inside the warehouse, every bad record gets stored first and cleaned after, which means you’re paying the destination’s meter to fix what the pipeline should have caught at the door. Garbage in, invoice out.

Two lanes

The way out is not abandoning the warehouse. Snowflake and Databricks are excellent at what they’re built for: analytics at scale, on enormous data, with the governance enterprises actually need. I’m not telling you to leave. I’m telling you to stop living there.

Picture two lanes.

Lane one is the lake: storage, analytical compute, the catalog, the dashboards. Deep, wide, and worth every penny when you use it for what it was built for.

Lane two runs beside it and handles ingestion and quality. Batch data comes in from the messy outside world (files, feeds, APIs, partner drops) and gets validated, transformed, and shaped before it lands anywhere. Lane two exists to deliver clean, trustworthy data to lane one. It does not exist to belong to lane one.

Because the moment lane one’s vendor owns lane two, you’ve lost the property that makes the whole architecture safe: you can no longer change your mind about the destination without rebuilding the road.

Respect, not allegiance

So hold lane two to a standard: full respect for the warehouses, zero allegiance to any of them.

Full respect means the pipeline treats each destination the way its own tools would. Point it at Snowflake and it should land proper, governed tables. Point it at Databricks and it should land managed Delta tables in Unity Catalog, real citizens of the lakehouse rather than files dumped over the fence for you to mop up.

Zero allegiance means the pipeline itself, the schema, the quality rules, the transformations, lives outside all of them. The same pipeline that feeds Snowflake should be able to feed Postgres for your application, object storage for retention, MongoDB when you need flexible access. One definition, validated once, delivered wherever the data needs to be.

When lane two clears that bar, questions that used to be migrations become configuration. Adding Databricks next to Snowflake is adding a destination. Trying a warehouse is pointing a pipeline at it. Leaving one (and someday, somewhere, you will want to leave one) is the same motion run backward. Your ingestion logic, your quality rules, everything your team knows about what clean data means in your business, none of it is hostage to the decision.

Data can have gravity. Logic shouldn’t.

One level down

When Databricks rebuilt its summit around governing agents and context, I wrote that they’d validated the idea of a governance layer, and that their version of it only fully works if you’re all in on their lakehouse. This is the same argument one level down the stack. It isn’t just the governance layer that should be neutral. It’s the lane that feeds the lake in the first place.

The vendors understand something worth taking seriously here: whoever owns the pipeline owns the default. Every record flows through it. Every quality decision happens inside it. If that layer is theirs, every future architecture question has a thumb on the scale.

Keep the thumb off the scale. Use the warehouses hard, for everything they’re great at, and keep the lane that feeds them yours.

Load the lake. Don’t live in it.


Todd Fearn is the founder and CEO of Datris, an open-source, agent-native data platform built on the Model Context Protocol. He also runs IData Corporation, the data engineering consultancy Datris grew out of. Before all this he spent thirty years building production data infrastructure inside financial institutions, among them Goldman Sachs, Bridgewater, Deutsche Bank, Salomon Brothers, and Freddie Mac. He has founded several venture-backed startups.