Can we run this ourselves? Yes. Here is the sizing.
Datris runs anywhere Docker runs. An evaluation fits on an 8 GB host with a 2 GB heap, and uploads and tap runs stream to disk, so memory does not grow with the size of the data passing through. Production wants more, and most of the stack is optional. Nobody gets budget approved without a table, so here is the table.
Self-host anywhere you want
Datris runs on-prem, in any cloud, or on your laptop — anywhere Docker runs. Open source under AGPL-3.0, built on proven infrastructure. No proprietary services, no vendor lock-in, no surprise bills. There is no managed service; your deployment lives inside your perimeter.
$ curl -fsSL https://get.datris.ai/install.sh | sh One command. Then open http://localhost:4200 in your browser. Your first pipeline in under a minute.
Planning a real deployment? See sizing and reference deployments.
Sizing
| Profile | Host | JVM heap | Notes |
|---|---|---|---|
| Evaluation / demo | 8 GB RAM, 4 vCPU, 40 GB disk | 2 GB | Postgres+pgvector, MongoDB, MinIO, Vault, ActiveMQ, MCP server, Assistant. Kafka, the extra vector stores, and the bundled embedding model off. |
| Recommended production | 16 GB RAM, 8 vCPU, 200 GB disk | 4 GB | Full stack on one node. Room for the bundled embedding model and concurrent taps. |
| High availability | 3 nodes × 16 GB RAM, 8 vCPU | 4 GB per node | Platform services on the nodes; Postgres, MongoDB, and object storage run externally (managed or your own clusters). |
Figures are starting points. The docs sizing page has per-component detail and tuning.
What is optional
| Component | When you need it |
|---|---|
| Apache Kafka | Only for streaming sources and Kafka destinations. An external Kafka needs no local broker. |
| PostgreSQL | Bundled by default as the structured destination and pgvector store. Turn it off, or point at a Postgres you already run. |
| Vector stores | Qdrant, Weaviate, Milvus, Chroma are optional; pgvector ships in the base Postgres. |
| Snowflake / Databricks | External accounts, not services Datris runs. Only if you land data there; credentials go in Vault. |
| Ollama | Only for local models in air-gapped deployments. Otherwise the AI features use a hosted provider key. |
| Bundled embedding model | The 2.2 GB bge-m3 model is needed only for air-gapped embeddings; otherwise use a hosted embedding provider. |
| Prometheus | Only if you want metrics scraped. The platform runs without it. |
Opt-in services are switched on with COMPOSE_PROFILES; the ones that ship enabled
(Postgres, the bundled embedding model) are switched off with a per-component flag. The
one-command install asks which you want.
Reference shape 1 · Single-node pilot
One host. Docker Compose. Postgres with pgvector, MinIO, MongoDB, Vault, ActiveMQ, the platform services, the MCP server, and the Assistant. Everything else off. This is the shape most teams evaluate on and many run in production for a single data team.
curl -fsSL https://get.datris.ai/install.sh | sh Then open http://localhost:4200 in your browser.
Reference shape 2 · HA production
Platform services on three nodes behind a load balancer. State moves out: managed or self-run Postgres, MongoDB, and S3-compatible object storage. Vault runs as your existing Vault cluster or a dedicated one. The optional components above, such as Kafka or a dedicated vector store, join only if the workload needs them.
Operating it
- Metrics. The server exposes a Prometheus endpoint for scraping, reachable only from inside the deployment's network.
- Audit to SIEM. Every audit entry is also written to the server log as one JSON line for your SIEM or log aggregator, and the log exports to CSV.
- Orchestration. An Airflow provider triggers Datris taps and monitors the resulting pipeline runs from your existing DAGs.
- Doctor. A built-in self-check catches the problems usually found after an outage: a Vault token about to expire, a missing AI provider secret, an embedding model never pulled, a disk filling up, a mixed-version stack. It runs at every start, on a timer to your incident webhook, before every upgrade in the installer, from the UI, the CLI, and an MCP tool. Every finding names the fix; nothing is changed automatically.
- Air-gapped. Local models via Ollama and the bundled embedding model; no telemetry and no required outbound calls. See Security.