One governed loop, every step recorded.
Datris is the open-source data control plane for AI agents. An agent asks for data; the platform acquires it, validates it, lands it in the stores you already run, and returns it with a receipt.
One agent-driven loop
Acquire, validate, land, observe, explain, repair. Same loop every run, same audit trail every time — so the agent's job is reasoning about the data, not improvising infrastructure.
Agents don't need a new data platform. They need a way into yours.
Datris sits beside the warehouse and the lake, never in front of them. It is the intake valve: the agent asks, Datris acquires and validates, and the rows land in the stores your teams already query. Nothing moves out of your stack and nothing new becomes the system of record.
Agents Don't Need a New Data Platform →
Load the Lake. Don't Live in It. →
- SnowflakeAnalytics warehouse
- DatabricksLakehouse
- PostgreSQLOperational store
- MongoDBDocuments
- S3 / MinIOObject storage, Parquet
- pgvector · Qdrant · Weaviate · Milvus · ChromaVector search
One pipeline lands the same validated records in several of these in parallel, idempotent by key. Keep the intake layer neutral and every future architecture decision stays yours.
Enforced by the platform, not by the prompt
The controls a risk committee asks about, in the open-source build, on by default. This is the part of Datris a head of data repeats to an auditor.
Not every answer needs a table
Most of what an agent asks for is a one-off: what a source returns for these ids, whether today's file is clean, what a transformation would produce. Live Read runs the full governed pipeline and hands the rows back to the caller instead of landing them. If the rows turn out to be worth keeping, switch the destination. The tap, its schedule, and its sync bookmark do not move.
- The tap, generated or hand-written, with credentials brokered through Vault
- Data quality rules and the transformation chain
- Agent Policy, the audit log, and the run record
- The agent never holds a key, in either lane
The built-in Assistant picks the lane for you: a one-off question goes to Live Read, not a new table. Live Read in the docs →
Your AI agents are
first-class pipeline operators
Datris ships with a native MCP server. Claude, Cursor, OpenClaw, and any MCP-compatible AI agent can register pipelines, trigger jobs, and query your structured, document, and vector data in real time — all through natural conversation.
- Register pipelines and generate schemas from sample data
- Create, schedule, and run AI-generated taps
- Ingest documents into vector databases (extract → chunk → embed)
- Upload data for processing
- Trigger and monitor pipeline jobs
- Profile data and get AI insights
- Semantic search across vector databases
- Query PostgreSQL, MongoDB, Snowflake, and Databricks directly
- Get an answer without landing a table — Live Read pipelines return validated rows straight to the agent
- Run the built-in doctor to diagnose the platform before retrying a failed run
- Manage credentials via Vault — without ever holding the key
Speaks every data language
Ingest structured data, unstructured documents, and archives. Output to vector stores, structured stores, or optimized columnar formats.
| Format | Input | Default Destination |
|---|---|---|
| CSV | SQL DB | |
| JSON / NDJSON | NoSQL DB | |
| XML | NoSQL DB | |
| Excel (.xlsx) | SQL DB | |
| Vector DB | ||
| Word (.docx) | Vector DB | |
| PowerPoint (.pptx) | Vector DB | |
| HTML | Vector DB | |
| Email (.eml) | Vector DB | |
| EPUB | Vector DB | |
| Archives (.zip, .tar, .gz) | Unpacked, routed | |
| Plain Text | Vector DB |
Destinations are fully configurable. Route any format to any target — SQL databases, NoSQL stores, vector databases, REST endpoints, Kafka topics, or ActiveMQ queues. Object-store destinations write Parquet, ORC, or Apache Iceberg tables with atomic commits, upserts by key, partitioning, and schema evolution. Or skip landing entirely: a Live Read pipeline runs the same validation and transformation chain and hands the rows straight back to the caller.
Full RAG pipeline built in
Extract, chunk, embed, and upsert documents into any major vector database. Build retrieval-augmented generation workflows without leaving your pipeline.