What it
costs

Data Oil is $0.02 per thousand units, with every rate behind it on your console before anything is charged. The comparison that matters is not our rate against somebody else's rate. It is one bill against the stack of bills you reconcile today.

The number

$0.02per 1,000 units

2 cents per thousand

A statement that creates, reads, updates or deletes is one unit. A search that ranks candidates is five to ten, by the data it examined.

Storage, media, documents, translation and the assistant sit beside it on one itemised card — 115 meters, every rate visible before and after it is charged.

What your work actually costs is checked against your budget at every stage, and work stops above 95% of it.

Every rate on the cardOpen the card
MeterRate, USDWhat it counts
Reading records$0.0145 per 1,000a statement that reads
Writing records$0.0199 per 1,000a statement that writes
Searching records$0.0725 to $0.145 per 1,000a search, five to ten units by the data it examined
storage.bytes$0.10 per GiBwhat your data occupies, for the period
storage.backup-bytes$0.05 per GiBwhat your backup chains occupy
store.egress-bytes$0.09 per GiBbytes sent out to you
media.audio-seconds$0.0667 per 1,000 saudio transcribed — $0.24 an hour
media.frames$1.00 per 1,000frames embedded — $0.06 a minute of video
document.pages$4.80 per 1,000pages, slides, sheets or sections extracted
translate.source-bytes$300.00 per GiBtext handed to the translator, once
translate.delivered-bytes$2,160.00 per GiBeach gigabyte delivered, so three languages is three deliveries
nlp.requests$0.0006 eachquestions to the assistant, plus $0.0005 and $0.02 per 1,000 tokens in and out
audit.records$0.006 per 1,000console audit lines
compute.instance-hour.*$0.184 to $4.328 an hourthe size you chose, held for you
compute.shared.processor-hour$0.060 a processor-houryour share of a shared node’s processors, minute by minute
compute.shared.memory-gb-hour$0.008 a GiB-houryour share of a shared node’s memory, minute by minute
compute.gpu$0.00the GPU — counted, and charged at nothing: transcription, frames, pages and translation already include it

Two small things you will see on your usage. Each time our ingestion service starts, it reads back each recent job's record from your own store — a few kilobytes a job, priced at $0.00 on the published card. And a submission of many files is charged file for file, as each would be sent alone, plus the statements that publish it: about $0.0012 on a 2,000-row table.

The size you run on

What a size contains

SmallMediumLargeXLarge
Processors and memory2 vCPU, 8 GiB4 vCPU, 16 GiB8 vCPU, 32 GiB16 vCPU, 64 GiB
GPUnoneshared, up to 8 GiB at a timeshared and ahead of others, up to 12 GiByours alone, 24 GiB
Ingestions at once1248
What runs on itrows and text+ documents, audio, video, translation+ the assistanteverything, shared with nobody
A month$134.32$268.64$537.28$3,159.44

A size is a reservation, charged for the hours it is held, with your usage on top. It bounds your connections, how many ingestions run at once, and how much GPU your work may hold — and that last one is what you feel: the same forty scanned pages took 278 seconds on Small and 133 on XLarge, measured on the running platform.

Dedicated or shared. Every size above is a node of your own. Or share a node with others and pay only your share of it, a minute at a time: $0.060 a processor-hour and $0.008 a GiB-hour of memory. On your own on a shared node, those rates come to exactly the Small, Medium or Large hour for the same shape — and your charges name the size you ran on and for how long. A shared node grows with demand and comes back down when it falls, always within its size’s range, and you pay for the size it actually ran at.

No size at all is a real choice. Run without one and you pay for what you use and nothing for capacity kept warm. It is the right answer until a workload needs the machine to be there before it starts, and it is where every customer on this platform began.

A bill you can check

500 GiB loaded, translated, and a hundred people asking

One real workload priced line by line at the card above: a 500 GiB CSV loaded once, a gigabyte of its free-text column translated into Japanese, and a hundred people asking the assistant about it for a month — twenty questions each a day, answered with citations.

$4,383.52the first month, everything in it
$3,490.58of that, paid once: the loading and the translation
$892.94every month after
The invoice, line by line500 GiB, one language, 100 users asking, 30 days
QuantityUSD
Ingestion, bytes taken in500 GiB$250.00
Ingestion, rows stored2.68 billion$1,342.18
Translation, source read1.00 GiB$300.00
Translation, Japanese delivered0.74 GiB$1,598.40
Queries, statements60,000 statements$0.87
Queries, rows returned1.2 million$4.80
Queries, bytes sent out2.9 GiB$0.26
Storage held500 GiB-month$50.00
Compute, the Large class held all month730 hours$537.28
Questions to the assistant60,000 questions$299.73
Total, first month$4,383.52
Every assumption behind itOpen the workings

Rows average 200 bytes, so 500 GiB is 2.68 billion rows. One gibibyte of the free-text column is translated into Japanese, which comes back at 0.74 times the source in bytes — measured, because an ideograph is three bytes and carries more than an English word. Each person asks 20 questions a day, the retrieval under each counted as one statement returning 20 rows and about 50 KB. The compute class is Large, held for the whole month.

Rows are loaded plain here. Knowledge extraction — entities and relationships in the graph — is a model call for every row and is charged only if you ask for it: $0.0015 per thousand, or $4,026.53 across this file.

Change any one of these and the arithmetic changes. Nothing here is a rounded figure, a plan or a minimum.

The loading and the translating are paid once. Month two is $892.94 — the storage, the machine and the asking. The hundred people asking come to $299.73, at $0.005 a question: each one rewritten, planned, retrieved four ways, reranked and answered with its citations, against your own model.

A search and a question are not the same thingWhy one is cheap and one is not

A retrieval product hands back passages and stops. A question here is rewritten, planned, retrieved four ways, reranked, answered and cited — and the timings are measured on our own hardware.

A searchA question to the assistant
What you give ita query — WHERE region = 'APAC' AND status = 'open', or the words you want matcheda sentence — “why did APAC renewals slip last quarter?”
What comes backthe rows that match, in the order you asked for. You read them and work out the answerthe answer, written out, with citations to the rows it used so you can check it
What actually runsone lookup against an indexthe question is rewritten and planned, retrieved four ways — vector, full text, graph, metadata — reranked by a cross-encoder, then a language model writes the answer
What it consumesmilliseconds of disk and index2.2 seconds of a GPU, measured
What a thousand of them cost$0.157.5 units each at the published price; $0.10 to $0.20 by the data examined$5.00half a cent each

Against the alternatives

The same workload, bought as three products

Everywhere else this workload is three products: a warehouse to hold it, a search product to make it findable, and a translation API. That is the comparison that matters, rather than any single vendor's list price.

The same workload, whole, either wayMonthly, USD
A warehouse, a search product and a translation APIBigQuery, Algolia and AWS Translate at list$1,358,653.17
Data Oil, one bill$4,083.79

333 times, for the same job. The assistant is out of both columns, because none of those three products answers a question at all.

The three products, and what each one listsRead the three lines
What you would buy insteadIts list, this workload
A warehouse to hold and query itBigQuery: storage, loading free, the same searches scanning a gigabyte each$369.76
A search product to make it findableAlgolia at $0.50 per 1,000 records held — every month, not once$1,342,177.28
A translation API for the JapaneseAWS Translate at $15 per million characters$16,106.13
All three, which is what this one workload needs$1,358,653.17

Any one of them is cheaper than us on its own, and none of them is the purchase you would actually make. A warehouse is cheaper because nothing is done to the data on the way in: not indexed for search, not embedded for meaning, not translated, not answerable in a sentence. The search line is the large one — 2.7 billion rows held at $0.50 per thousand, every month, forever. The real alternative is to index part of your content and accept that the rest is not searchable. Here everything is indexed on the way in, once, inside the $1,592.18 ingest line.

Every rate on the card, against the product it replaces14 rates against 14 vendors
What you are buyingData OilWhat you would buy insteadOurs, as a share
Reading a thousand records$0.0185 per 1,000Firestore, $0.03 per 100,000 document reads$0.0003 per 1,0006167%
Writing a thousand records$0.0239 per 1,000Firestore, $0.09 per 100,000 document writes$0.0009 per 1,0002656%
A search$0.15 per 1,000Algolia, $0.50 per 1,000 searches, and $0.40 per 1,000 records held every month$0.5030%
A row made retrievable$0.0005 per 1,000Pinecone, $4.00 to $4.50 per million write units$0.004 per 1,00012%
A gigabyte taken in$0.50 per GiBturbopuffer, $2.00 per GB written$2.1523%
A gigabyte held$0.10 per GiB-monthNeon $0.35, turbopuffer $0.33, Supabase $0.125$0.3827%
A gigabyte of backup held$0.05 per GiB-monthAWS RDS backup beyond the free allowance, $0.095 per GB-month$0.1049%
A gigabyte sent out$0.09 per GiBAWS, Supabase and Typesense all at $0.09 per GB$0.1093%
An hour of processor$0.20 per hourNeon, $0.42 per vCPU-hour$0.4247%
An hour of audio transcribed$0.24 an hourDeepgram $0.26, Whisper $0.36, Google $0.96, AWS $1.44$0.2693%
A minute of video$0.06 a minuteRekognition and Video AI, $0.10 a minute$0.1060%
A thousand document pages$4.80 per 1,000Azure and Google layout parsers, $10 per 1,000$10.0048%
A gigabyte translated, one language$1,898.40 per GiBAWS Translate, $15 per million characters$16,106.1312%
A question to the assistant$0.005 a questionVertex AI Search, $4 per 1,000 queries — retrieval only, before any answer is written$0.004 a query125%

Vendor list prices, each converted into the unit this card meters in. Reading and writing include the row returned or written; a search is priced at 7.5 units of the published $0.02, the middle of its band. Egress is a pass-through — everyone resells the same bandwidth at the same price. Transcription is under all four of the transcription APIs we compared. The two Firestore rows go the other way and are on the table for that reason: Firestore charges per document read and is cheaper at it than we are, because a key-value read is the one thing it is built to sell — and it ranks no text at all, so Google's own documentation tells you to buy a search engine beside it. The assistant row is the one where a percentage misleads: $4 per thousand queries elsewhere buys retrieval and stops there, while a question here is rewritten, planned, retrieved four ways, reranked, answered and cited for $0.005.

Two more shapes

A knowledge base, and a product catalogue

A knowledge base

50,000 documents and 100,000 searches a month. Rows, full-text, vectors and storage — the same work on both sides.

Supabase Pro with Algolia beside it for search: $74.37.

Data Oil, all of it on one bill: $24.45 — about a third, with the vectors and the graph already in it.

A product catalogue with search

250,000 records, 2,000,000 reads and 500,000 searches a month.

Firestore with Algolia, which is Google's own guidance for full-text search: $307.97.

Data Oil, everything in one: $123.80.

A million searches a month on a bolt-on search vendor is $495.00. The same million on Data Oil is $150.00.

Translation

Priced by the gigabyte, and cheaper as you use more

Delivered in a monthPer GiB of sourcePer GiB delivered, each language
The first 10 GiB$300.00$2,160.00
10 to 100 GiB$240.0020% less$1,728.0020% less
Beyond 100 GiB$180.0040% less$1,296.0040% less

One gibibyte into one language is $1,898.40. Fifty gibibytes into one language is $80,856.00, where AWS Translate lists the same fifty at $805,306.37 — 10x more. Every further language adds only the text delivered in it: the source read is charged once however many languages you ask for, where a translation API charges again, in full, for each one.

Translating everything is a project, and we plan it with you

Translating at that scale is bound by time before it is bound by money. Tell us the volume, the languages and the date you need it by, and we come back with the options — what each one costs, how long each one takes, and which of them fits the budget you have. You choose before anything starts. Translation is shared across every translation server your deployment has, so more at once is more servers, each with its own GPU.

Plan a large translation with us

Or have it fixed

Send three months of invoices. Get a fixed price.

Send the invoices from your database, search, vector and AI vendors for the last three months. You receive a fixed monthly price for the same workload on Data Oil, held for twelve months, with the line-by-line workings.

The invoices are read by one person, and are returned or deleted once the analysis is sent.

A fixed price is a number finance can put in a budget before anything is signed.

Measure it, or have it fixed.

A proof of concept leaves you with the forecast measured on your own data. The invoices leave you with a fixed price held for a year. Either is a number you can take to finance.

Talk to us