Nobody designed the stack you run. It accumulated.
A database. Then a search engine beside it, because the database could not rank. Then a vector index, because search could not understand. Then a translation service, an AI vendor, and a sync job between every pair. We ran that stack. Data Oil is what we built to stop running it.
Five short chapters. No feature table until you ask for one.

Chapter one
We understand where you are
These are the sentences we hear first, in the words we hear them. If two of them are yours, the rest of this page was written for you.
Search is a second system, and it is never quite in step with the database.
A sync job, a queue, and a Monday spent finding out which records the index does not have.Every AI project starts with an export.
The model cannot see the data where it lives, so somebody copies it out, and the copy is stale by the demo.Nobody can tell me what it costs until the invoices arrive.
Six vendors, six units, six bills, and a spreadsheet that reconciles them a month late.Every new kind of data is a new vendor.
Audio, video, PDFs, a graph: each one a procurement, a credential, a contract and another sync.Our customers work in eighteen languages and our search works in one.
A translation pipeline that runs after the fact, and a search index that only ever saw the original.The security review is where we lose weeks.
Who can see which column, who saw it, and what exactly leaves for the model provider - answered in slides, never in a demonstration.
None of these is a failure of your team. Each is what happens when a demand arrives that the stack was not shaped for, and the only move available is to bolt on one more system.
Chapter two
Why it happened, and why it keeps happening
Every one of those systems was designed for one shape of data. The demands did not stay in one shape.
- Then
Rows in a database. Search was a LIKE clause and nobody minded.
- Then
Search had to rank, so a search engine arrived beside the database, with a job to keep them in step.
- Then
Search had to understand meaning, so a vector index arrived beside both.
- Then
The data had to be asked questions, so a model arrived, fed by exports.
- Then
Users in other languages, documents, audio, video, a graph of who is connected to what.
- Now
Agents that act on the data, with credentials and budgets of their own, and auditors who want every act.
The stack did not fail you. It ran out of shapes. Each new demand became a new system, and each new system became a seam - and the seams are where the time, the money and the incidents go.
What follows is the alternative: one engine that stores every shape, indexes every shape, and answers over all of them in one statement.Chapter three
What changes with Data Oil
Take the six sentences from chapter one, and a seventh we hear from the people who build. Here is what each becomes.
Search is a second system.
what we heardIt is the same system.
A record commits once and is searchable by keyword, by meaning and by relationship the moment it does. One engine builds every index over the same records - no second query language, and nothing to keep in step.
How it works
One statement filters on a column, ranks by vector distance and scores a keyword match:WHERE SEARCH_INDEX(...) AND filedAt > ... ORDER BY VECTOR_DISTANCE(...).Every AI project starts with an export.
what we heardThe model sees the data where it lives.
Ask in plain language and the answer comes back with the evidence behind it, retrieved from the database itself, against the model endpoint you choose - yours, with your key. Nothing is copied out to be asked about.
How it works
And before anything leaves:POST /query/preparehands back the exact submission that would have gone to the model, without sending it. Your security team can read it.Nobody can tell me what it costs.
what we heardThe cost is on the console before it is charged.
One published unit - 2 cents per thousand - and every rate behind it on your console. What work actually costs is checked against the budget you set at every stage, and work stops above 95% of it.
How it works
Every rate is visible to you on the console before and after it is charged, and the bill is the same figures.Every new kind of data is a new vendor.
what we heardA new kind of data is a meter, not a procurement.
Files, tables, images, audio, video, PDFs, Office documents, e-mail, web pages - thirty-one formats, read and understood on the way in and stored beside the original. A document's tables become rows you can query. Send a hundred files at once, or a zip of them, and they arrive together.
How it works
The audio you transcribe and the pages you extract appear on the same bill as the rows you read.Our search works in one language.
what we heardWritten once, searchable in eighteen.
Text is translated on write into up to eighteen languages and indexed in each, so a search in one language finds a document written in another - and the answer comes back in the reader's.
How it works
Translation is priced by the gigabyte of source read once, and by the gigabyte delivered in each language - because each language is a full pass over the whole text, which is where the work is.Every integration starts with an export, and there is no SDK.
what we heardA package in your language, a retriever in your framework, a server for your agent.
Packages for .NET, Python and Node, retrievers for LangChain and LlamaIndex that ask the database for passages instead of exporting it, and an MCP server that gives Claude Desktop, Cursor or any agent four tools over your data.
How it works
A retrieval is a statement, so every fence the engine applies to a statement applies to what the model is handed.The security review is where we lose weeks.
what we heardIt is where this platform is strongest, so it goes first.
Rights down to a column, the same fence on every door, residency per database that refuses a write outside its region, and an audit line for every act - including ours.
How it works
In the proof of concept the security demonstrations run in week one, live, in the order your reviewers will ask for them.
Chapter four
Today, and the demands still coming
The next demand is already visible. The question is whether it will be a meter or another migration.
What is arriving now
- Agents that act on your data with a credential and a budget of their own, and are refused in a way they can read.
- Questions that have to be answered with citations, not summaries.
- Documents by the page, media by the minute, tables inside PDFs.
- Data that has to stay in a region, per customer, provably.
- A model provider you may want to change next year.
How Data Oil meets it
- An API key scoped to a database and four verbs, a quota it cannot exceed, and every act it took in your audit.
- Answers carry the evidence they came from, and the assistant is a route on the same surface as the rows.
document.pages,media.audio-seconds,media.frames: meters that already exist.- Placement and residency decided per database; a write outside the region stores nothing.
- The model endpoint is a setting per database, with your key, changed without a migration.
Moving to Data Oil is not the last migration you will make because we say so. It is the last one because the next demand lands as a meter on a platform that already stores every shape.
In production today behind three products, each on its own infrastructure; the same engine at every size.On the calendar
Three things you can plan against
- September 2026
PostgreSQL, against the tools you already own. DataGrip, DBeaver, Power BI and the common drivers, each measured tool by tool. Done.
- September 2026
An Entity Framework Core provider. LINQ straight to the engine, through the same door as everything else. Done.
- Now
The packages. .NET, Python, Node and MCP, built and tested against the reference deployment.
Chapter five
Moving is a serious decision. We make it a measured one.
Nobody should move a system they depend on because a website was persuasive. So the first step is not a contract. It is a written set of goals, on your data, with a date.
- The one workload the proof has to carry - a multilingual search, a set of documents that has to answer with citations, a reporting store replacing two systems, a migration off a database-plus-search-engine pair.
- Twenty questions it has to answer, written by the person on your side who knows what the answers should be.
- The measures: relevance on the twenty, the queries the reporting has to run, the languages search has to work in, the load that has to complete, the restore that has to be proved.
- The date: four weeks from the day the data lands. And a written answer at the end, whichever way it goes.
What you risk
Four weeks of one person's attention and a fixed fee. The fee is credited at 10% of your recurring charges each month for the first year, up to the fee itself - against storage, compute, statements, searches and the assistant, and never against ingestion and translation. You know the fee when you sign, so you know the cap. A buyer who does not proceed paid for a report they can act on.
What you leave with
A deployment you signed into, your real data indexed four ways, the query you cannot run today measured against the one you run now, a restore you watched, and the metered usage accruing on the console from day one - which is the forecast for production, itemised by meter.
The first step
Tell us the three things that hurt most.
Not a demo request. Three sentences about what you run and what it costs you in time, in money or in sleep. We reply with whether Data Oil is the answer - and if it is not, we say so, because a proof of concept that will fail is a month of your life we would rather not spend.