Chapter one

We understand where you are

These are the sentences we hear first, in the words we hear them. If two of them are yours, the rest of this page was written for you.

  • Search is a second system, and it is never quite in step with the database.A sync job, a queue, and a Monday spent finding out which records the index does not have.
  • Every AI project starts with an export.The model cannot see the data where it lives, so somebody copies it out, and the copy is stale by the demo.
  • Nobody can tell me what it costs until the invoices arrive.Six vendors, six units, six bills, and a spreadsheet that reconciles them a month late.
  • Every new kind of data is a new vendor.Audio, video, PDFs, a graph: each one a procurement, a credential, a contract and another sync.
  • Our customers work in eighteen languages and our search works in one.A translation pipeline that runs after the fact, and a search index that only ever saw the original.
  • The security review is where we lose weeks.Who can see which column, who saw it, and what exactly leaves for the model provider - answered in slides, never in a demonstration.

None of these is a failure of your team. Each is what happens when a demand arrives that the stack was not shaped for, and the only move available is to bolt on one more system.

Chapter two

Why it happened, and why it keeps happening

Every one of those systems was designed for one shape of data. The demands did not stay in one shape.

  • Then

    Rows in a database. Search was a LIKE clause and nobody minded.

  • Then

    Search had to rank, so a search engine arrived beside the database, with a job to keep them in step.

  • Then

    Search had to understand meaning, so a vector index arrived beside both.

  • Then

    The data had to be asked questions, so a model arrived, fed by exports.

  • Then

    Users in other languages, documents, audio, video, a graph of who is connected to what.

The stack did not fail you. It ran out of shapes. Each new demand became a new system, and each new system became a seam - and the seams are where the time, the money and the incidents go.

What follows is the alternative: one engine that stores every shape, indexes every shape, and answers over all of them in one statement.

Chapter three

What changes with Data Oil

Take the six sentences from chapter one, and a seventh we hear from the people who build. Here is what each becomes.

  • Search is a second system.what we heard

    It is the same system.

    A record commits once and is searchable by keyword, by meaning and by relationship the moment it does. One engine builds every index over the same records - no second query language, and nothing to keep in step.

    How it worksOne statement filters on a column, ranks by vector distance and scores a keyword match: WHERE SEARCH_INDEX(...) AND filedAt > ... ORDER BY VECTOR_DISTANCE(...).
  • Every AI project starts with an export.what we heard

    The model sees the data where it lives.

    Ask in plain language and the answer comes back with the evidence behind it, retrieved from the database itself, against the model endpoint you choose - yours, with your key. Nothing is copied out to be asked about.

    How it worksAnd before anything leaves: POST /query/prepare hands back the exact submission that would have gone to the model, without sending it. Your security team can read it.
  • Nobody can tell me what it costs.what we heard

    The cost is on the console before it is charged.

    One published unit - 2 cents per thousand - and every rate behind it on your console. What work actually costs is checked against the budget you set at every stage, and work stops above 95% of it.

    How it worksEvery rate is visible to you on the console before and after it is charged, and the bill is the same figures.
  • Every new kind of data is a new vendor.what we heard

    A new kind of data is a meter, not a procurement.

    Files, tables, images, audio, video, PDFs, Office documents, e-mail, web pages - thirty-one formats, read and understood on the way in and stored beside the original. A document's tables become rows you can query. Send a hundred files at once, or a zip of them, and they arrive together.

    How it worksThe audio you transcribe and the pages you extract appear on the same bill as the rows you read.
  • Our search works in one language.what we heard

    Written once, searchable in eighteen.

    Text is translated on write into up to eighteen languages and indexed in each, so a search in one language finds a document written in another - and the answer comes back in the reader's.

    How it worksTranslation is priced by the gigabyte of source read once, and by the gigabyte delivered in each language - because each language is a full pass over the whole text, which is where the work is.
  • Every integration starts with an export, and there is no SDK.what we heard

    A package in your language, a retriever in your framework, a server for your agent.

    Packages for .NET, Python and Node, retrievers for LangChain and LlamaIndex that ask the database for passages instead of exporting it, and an MCP server that gives Claude Desktop, Cursor or any agent four tools over your data.

    How it worksA retrieval is a statement, so every fence the engine applies to a statement applies to what the model is handed.
  • The security review is where we lose weeks.what we heard

    It is where this platform is strongest, so it goes first.

    Rights down to a column, the same fence on every door, residency per database that refuses a write outside its region, and an audit line for every act - including ours.

    How it worksIn the proof of concept the security demonstrations run in week one, live, in the order your reviewers will ask for them.
Arrivesa file or tablea recordingRead31 formatsOCR for scansChunkedsectionstables to rowsUnderstoodembeddingsentitiesTranslatedup to eighteenlanguagesSearchablekeywordmeaningthe original is kept beside every derived form, every stage held to your budget
What arrives is read, chunked, understood and translated on the way in, and the original stays beside every derived form.

Chapter four

Today, and the demands still coming

The next demand is already visible. The question is whether it will be a meter or another migration.

What is arriving now

  • Agents that act on your data with a credential and a budget of their own, and are refused in a way they can read.
  • Questions that have to be answered with citations, not summaries.
  • Documents by the page, media by the minute, tables inside PDFs.
  • Data that has to stay in a region, per customer, provably.
  • A model provider you may want to change next year.

How Data Oil meets it

  • An API key scoped to a database and four verbs, a quota it cannot exceed, and every act it took in your audit.
  • Answers carry the evidence they came from, and the assistant is a route on the same surface as the rows.
  • document.pages, media.audio-seconds, media.frames: meters that already exist.
  • Placement and residency decided per database; a write outside the region stores nothing.
  • The model endpoint is a setting per database, with your key, changed without a migration.

Moving to Data Oil is not the last migration you will make because we say so. It is the last one because the next demand lands as a meter on a platform that already stores every shape.

In production today behind three products, each on its own infrastructure; the same engine at every size.

On the calendar

Three things you can plan against

  • September 2026

    PostgreSQL, against the tools you already own. DataGrip, DBeaver, Power BI and the common drivers, each measured tool by tool. Done.

  • September 2026

    An Entity Framework Core provider. LINQ straight to the engine, through the same door as everything else. Done.

The full roadmap, and what is under consideration.

Chapter five

Moving is a serious decision. We make it a measured one.

Nobody should move a system they depend on because a website was persuasive. So the first step is not a contract. It is a written set of goals, on your data, with a date.

What is agreed in writing in week one, before anything else
  1. The one workload the proof has to carry - a multilingual search, a set of documents that has to answer with citations, a reporting store replacing two systems, a migration off a database-plus-search-engine pair.
  2. Twenty questions it has to answer, written by the person on your side who knows what the answers should be.
  3. The measures: relevance on the twenty, the queries the reporting has to run, the languages search has to work in, the load that has to complete, the restore that has to be proved.
  4. The date: four weeks from the day the data lands. And a written answer at the end, whichever way it goes.

What you risk

Four weeks of one person's attention and a fixed fee. The fee is credited at 10% of your recurring charges each month for the first year, up to the fee itself - against storage, compute, statements, searches and the assistant, and never against ingestion and translation. You know the fee when you sign, so you know the cap. A buyer who does not proceed paid for a report they can act on.

What you leave with

A deployment you signed into, your real data indexed four ways, the query you cannot run today measured against the one you run now, a restore you watched, and the metered usage accruing on the console from day one - which is the forecast for production, itemised by meter.

The first step

Tell us the three things that hurt most.

Not a demo request. Three sentences about what you run and what it costs you in time, in money or in sleep. We reply with whether Data Oil is the answer - and if it is not, we say so, because a proof of concept that will fail is a month of your life we would rather not spend.