You write one statement,
not three.
The alternative is a database, a search engine and a vector store, with your code holding them together. Every schema change is three schema changes, every backfill is three backfills, and every incident starts with working out which of the three is behind. Data Oil is one SQL surface over all of it.
The statement
Rows, vectors, keywords and relationships in one query
SELECT title,
VECTOR_DISTANCE(embedding, EMBED('casing failure'), 'cosine') AS distance,
$score AS keywordScore
FROM Report
WHERE SEARCH_INDEX('Report[body]', 'casing failure') = true
AND filedAt > date('2026-01-01')
ORDER BY distance
LIMIT 10;
One index build, one transaction, one bill. A write commits once and is searchable by keyword, by meaning and by relationship the moment it commits.
Parameters carry their type
:price passed as a number compares numerically; a date stays a date.
Answers stream
Rows arrive one at a time, with a cursor that is a held enumerator rather than an offset - so paging through a database that is being written to never skips or repeats a row.
Errors are structured
RFC 9457 problem documents, a stable errorCode, and a retryable header - so a client decides what to do before it parses anything.
The packages
Install it, rather than write it
Sixty lines gets you a working client. The packages add what those sixty lines leave out: the headers, the streaming, typed errors and the retries.
DataOil.Client · NuGet
dotnet add package DataOil.Client. .NET 8 and 10, no dependencies beyond the framework. Streams the NDJSON, follows the cursor, types every refusal, retries a lost race.
dataoil · PyPI
pip install dataoil. Standard library only. The same client, plus retrievers for LangChain and LlamaIndex with dataoil[langchain] and dataoil[llamaindex].
@dataoil/client · npm
npm install @dataoil/client. JavaScript and TypeScript on fetch, no dependencies; query() is an async iterator over the streamed body.
@dataoil/mcp · npm
npm install -g @dataoil/mcp. A Model Context Protocol server over one database - query, schema, ask, ingest - for Claude Desktop, Cursor and any MCP client, on the key's own rights, budget and audit.
Scheduled
The next three months
| Scheduled | What lands | When |
|---|---|---|
| PostgreSQL tools | The PostgreSQL port measured against the tools people already own - DataGrip, DBeaver, Power BI, psycopg, Npgsql and the JDBC driver - each one connecting, browsing and querying, with a compatibility table per tool: connects, browses, queries, parameters, COPY, and what is refused and why. | shipped in September 2026 |
| Entity Framework Core provider | DataOil.EntityFrameworkCore, built on the NuGet client: LINQ translated to the engine's own dialect and sent through /query, so it inherits every fence, meter and audit line; your model checked against each database's schema as it is today; untranslatable queries refused rather than run on your machine; SaveChanges guarded by each record's version. | shipped in September 2026 |
| The packages | DataOil.Client, dataoil, @dataoil/client and @dataoil/mcp, built and tested against the reference deployment, are delivered with a deployment today - from Data Oil's own package directory, not the public registries - and go to nuget.org, PyPI and npm when the public release is decided. | delivered now; public registries to be decided |
+Four ways in, a client under sixty lines, and everything a store should haveOpen
Four ways in
HTTP /api/v2, embedded in a .NET process with no socket at all, the PostgreSQL wire protocol for any driver or BI tool, and a command line. The same engine and the same statements.
A client under sixty lines, and a package
In C#, Python or JavaScript, and the reference manual shows every route and every statement in all three, generated from one source so they cannot drift.
Everything a store should have
Records by identifier, transactions as resources, batches as one unit, and a change feed over a WebSocket with a bookmark you resume from.
Tell it what your data means
A description is schema, and the assistant reads it
CREATE PROPERTY Order.placedAt DATETIME (DESCRIPTION 'When the customer placed the order, in the shop''s local time.');
What the description does
The assistant reads it. "Which orders were late last quarter" resolves against your words rather than a guess from a column name. A description is data, never an instruction.
Then ask it a question
POST /query
X-Job-ID: ask-1041
X-Database: acme_corp
{ "query": "which suppliers missed a delivery window in Q3",
"collection": "deliveries", "max_results": 20 }The answer carries its sources, its citations, a confidence score and per-stage timings. And POST /query/prepare hands back the exact submission that would have gone to your model - prompt, evidence and byte count - without sending it. Let your security team read it.
What is in the box
Built from scratch for object storage, and it shows
A storage engine, a vector index and a full-text index all written from scratch for object storage, and a pipeline that transcribes, embeds, extracts and translates. 534 named reliability guarantees run on every build.
+The engine, the pipeline, ingestion and operationsOpen the box
The engine
A log-structured storage engine written for object storage. A SQL dialect with joins, window functions, common table expressions, set operations, graph traversal and pattern matching; an HNSW vector index and a BM25 full-text index written from scratch; snapshot isolation with optimistic commit; a change feed off the write-ahead log; and a PostgreSQL wire-protocol endpoint that speaks the extended protocol.
The pipeline
Transcription, image understanding, embedding, tabular parsing and translation into eighteen languages; entity and relationship extraction into the graph; a metadata and lineage service; and a query orchestrator that rewrites, plans, retrieves across vector, full-text, graph and metadata, ranks, and answers with citations. Nothing caps what you send: a file of any size streams into your own bucket and is processed where it lies.
Ingestion held to your budget
Three doors: a file, a staged upload of any size into your own bucket, a fetched URL. Every job is metered at your rates and checked against your budget at every stage, stopping above 95% of it; checkpointed jobs resume from durable partial results; a restart route re-runs from a stage.
Operations you can script
dataoil doctor exits 2 when anything is critical. dataoil schema diff a b exits 2 when they differ. EXPLAIN and PROFILE on any statement; a slow-query log grouped by statement shape; four lag measures per database.
Start now
Three calls to your first row
A credential and a database name
On your welcome message. see the manuals for further details takes you from a credential to rows coming back in ten minutes.
A statement
CREATE DOCUMENT TYPE,INSERT,SELECT. Or your PostgreSQL driver, connected to the wire port, with no adapter.A question
POST /querywith a sentence, and an answer with citations against the model you chose. Bring your own model is a configuration line.
Bring your workload
A paid proof of concept on your own data, with you at the keyboard
Four weeks, one workload, your data, your model endpoint, weekly working sessions with the people who built the engine. You leave with a running tenancy, a schema export, the statements, the measured comparison and the report, whichever way it goes.
Or send the invoices from your current stack and get a fixed price for the same workload on Data Oil.
Questions this page raises
Three for your next retrospective
How many places does a schema change land?
If the answer is three, every backfill is three backfills.
Read the answer →Does your pagination skip a row under writes?
An offset does. A held enumerator cannot.
Read the answer →When did the docs last run?
Every example here is parsed by the build before it is published.
Read the answer →