Questions worth asking

Seventeen questions.
Ask them of what you run today. Then ask us.

Every fact on this site implies a question about your own estate. Here they are, with what to ask your current vendors and the answer Data Oil gives. If yours is not here, send it; it is answered by the people who built the platform. Below the questions, a comparison of Data Oil beside Databricks, Snowflake and Amazon Redshift.

How many bills does one search feature generate today?

Ask this of what you run today

Count the vendors that hold a copy of the data your search reads: the database, the search engine, the vector store, and the model provider on the end. Each is a bill, a credential set and a job that keeps it in step with the others.

Data Oil's answer

One. Data Oil is one SQL surface over CRUD, vector, full-text and multilingual data, so the search reads the same records the transaction wrote, and there is nothing to keep in step. The pricing page shows the three-vendor stack beside the one bill, and the invoice analysis maps your actual bills into it.

What leaves the building when somebody asks an AI a question?

Ask this of what you run today

Ask your current vendor for the bytes: the system prompt, the evidence, the citations, the byte count, for one question one user asked. Ask whether the answer came from their model or yours, and where the key lives.

Data Oil's answer

Data Oil loads no model for the final answer. Your endpoint and key are registered, sealed and assigned per database. POST /query/prepare runs every stage inside the deployment and returns the exact submission, prompt, evidence and byte count, without sending it. Your security team can read it before anything leaves.

When did your backup last restore, not run?

Ask this of what you run today

A backup that completed is a file. A backup that restored is a proof. Ask when the last restore happened, who watched it, and what it restored into.

Data Oil's answer

Every Data Oil chain is proved by a real restore drill into a scratch name on the cadence you choose, daily, weekly or monthly, and the console shows proved and the date per database. You restore your own copy from the console, into a new name, whenever you like.

What stops the next bulk load running past its budget?

Ask this of what you run today

Most platforms tell you afterwards. Ask what the last one cost, how you found out, and whether anyone could have stopped it.

Data Oil's answer

Data Oil checks what every job has actually spent against the budget you set, at every stage of the work, and stops it above 95% of that budget. Every charge accrues live at your own rates, itemised by meter, so the bill holds no surprises.

Can finance reconcile the data bill to a meter, and to the team that incurred it?

Ask this of what you run today

Take last month's invoice and try to attribute one line to one service. If it takes a spreadsheet and a guess, that is the question.

Data Oil's answer

115 meters, every rate visible in the UI before and after it is charged, dated charge records kept 400 days that add up to the figure on the usage page, and every operation readable per request keyed on the job identifier you chose. A bill attributes to a service, not to a guess.

Who owns the job that keeps the search index in step with the rows?

Ask this of what you run today

There is a function, a queue or a cron somewhere that re-indexes on every write. Ask who owns it, what happens when it is behind, and how a schema change reaches all three stores.

Data Oil's answer

Nobody, because there is no job. Every index is built by the same engine over the same records inside the same transaction. A write commits once and is searchable by keyword, by meaning and by relationship the moment it commits. A schema change is one schema change.

How does your current store decide who the writer is?

Ask this of what you run today

Two servers that both believe they are the writer is the failure that loses data. Ask how yours decides, and what happens during a network partition.

Data Oil's answer

The same single-writer model as PostgreSQL, with one difference: the lease is arbitrated by the object store, not by a process that can lose contact with its replica. Two servers cannot both hold the lease, so there is no split brain by construction. A replacement takes over from the store, the promotion is an explicit, audited act, and every durable byte is already in the store, so a node that dies takes nothing with it.

What stops a write landing in the wrong region?

Ask this of what you run today

A policy document records where data should be. Ask what refuses a write that goes somewhere else.

Data Oil's answer

Residency is enforced per write against the region of the instance the database sits on. A write outside the declared region is refused and nothing is stored. You can hold databases on several providers and regions under one tenancy, and each one is fenced by its own placement.

What happens to the price when you outgrow the plan?

Ask this of what you run today

A fixed plan is cheap until the month it is not. Ask what the next tier costs, and what a cluster costs when it is idle.

Data Oil's answer

Usage pricing has no plan to outgrow. Two cents per thousand units from the first request, storage at $0.10 per GiB-month, and no cluster to size. The fixed-price offer holds a number for twelve months at the volume your invoices described.

Which vendor will do the second language?

Ask this of what you run today

Search that works in one language usually means a second product for the second language, and a third for the translation. Ask who does it and what it costs.

Data Oil's answer

The one you already have. Name the languages on the ingest and the content is translated on write into up to eighteen languages, stored beside the original and searchable in each, at $300.00 a gigabyte of source text. A language that cannot be delivered fails the job rather than being quietly omitted.

Does your pagination skip a row under writes?

Ask this of what you run today

An offset over a table that is being written to skips or repeats. Ask what the cursor actually is.

Data Oil's answer

A cursor on Data Oil is a held enumerator, not an offset, so paging through a database that is being written to never skips or repeats a row. Rows stream as NDJSON, a page at a time, and closing the connection cancels the query.

When did the documentation last run?

Ask this of what you run today

Documentation drifts from the product the day after it is written, unless something runs it.

Data Oil's answer

Every SQL example in the Data Oil reference manual is parsed by the engine's own build before it is published, and every route and statement is shown in C#, Python and JavaScript generated from one source so the three cannot drift. 534 named reliability guarantees run on every build of the engine.

What are the twenty questions?

Ask this of what you run today

The ones your users ask every day, the ones your current system cannot answer, and the ones your security team will ask about the answers.

Data Oil's answer

They are what the paid POC is built around. You provide them in a list, the success criteria are written against them in week one, and the report at the end says how each was answered, with citations you can check, against the model you chose.

What is the number that decides it?

Ask this of what you run today

A proof of concept without a number ends in a conversation. Pick one: latency on a query, cost on a month, or a query that cannot be run today.

Data Oil's answer

It is agreed in writing in week one and tested rather than argued about. The console shows the metered usage accruing from day one, so the cost number is measured, not modelled, and the report at the end says what it came to.

What can your agent reach that you did not grant?

Ask this of what you run today

An agent that holds the application's credential holds everything the application can do. Ask what the agent's key is scoped to, who can revoke it, and what happens to the one you cannot find.

Data Oil's answer

An agent on Data Oil holds its own key, minted with POST /api/v2/tokens for named databases and a permission, with an expiry, listed without its secret, and revoked in one call; a key may not mint a key. Roles fence what it may read down to the column on both query surfaces, and a quota refuses it before a statement is parsed, with the arithmetic in the refusal.

Can you show what your agent sent to the model last Tuesday?

Ask this of what you run today

Every question an agent asks becomes a submission to a model provider. Ask for the bytes of one of them, and for the audit line that says which agent asked and when.

Data Oil's answer

POST /query/prepare returns the exact submission, prompt, evidence and byte count, without sending it, so you can read what an agent would send before it does. Every statement and every question carries the job identifier the agent chose, and lands in your audit line by line; nothing on the customer surface can remove one.

Where does the data have to be?

Ask this of what you run today

Some data may not leave a region, a provider or a building. Ask whether the platform runs there, or only where the vendor is.

Data Oil's answer

Both. Data Oil runs as a hosted service on infrastructure we operate, or in your own cloud account against your own buckets with your own keys, and the console brings a node up on a machine you name. A POC in your environment is week one.

Comparison

Data Oil beside Databricks, Snowflake and Amazon Redshift

Ten dimensions, side by side, for the reader who wants the platforms on one page before asking any of the questions above.

Feature / DimensionData OilDatabricksSnowflakeAmazon Redshift
Core PhilosophyNLP Emotionally intelligent analyticsUnified data lake house for AI/ML and BICloud-native data warehousing with elastic scalingTraditional cloud data warehouse on AWS
Data ArchitecturePersona-aware pipelines, modular ingestionLakehouse: combines data lakes and warehousesMulti-cluster shared data architectureColumnar storage, MPP architecture
Ease of UseDesigned for clarity and emotional resonancePowerful but complex for non-technical usersIntuitive SQL-first interfaceFamiliar to AWS users, but less intuitive overall
ML & AI IntegrationBuilt-in emotional intelligence layerDeep ML/AI support via notebooks and MLflowLimited native ML; relies on integrationsBasic ML via SageMaker integration
Pricing ModelTransparent, consumption pricingUsage-based, can be costly at scaleConsumption-based, with auto-scalingInstance-based, less flexible
Deployment FlexibilityCloud-agnostic, hybridMulti-cloud (Azure, AWS, GCP)Multi-cloud (AWS, Azure, GCP)AWS-only
Differentiation StrategyEmotional clarity, regional trust, modularity, MultilingualPerformance and scale for data scienceElasticity, simplicity, and ecosystem integrationsTight AWS integration, legacy familiarity
Investor NarrativeGrounded in AI innovation with global ambitionOpen-source credibilityEnterprise-ready, IPO-backed growthTrusted AWS brand, conservative scaling
Community & EcosystemEmerging, Multi faceted dataStrong open-source and enterprise communityGrowing enterprise ecosystemMature AWS ecosystem
Use Case Sweet SpotRegional/Advanced analytics, AI, NLP and no codeAdvanced analytics, ML workflowsBI dashboards, scalable warehousingTraditional BI, ETL-heavy workloads

Your question

Send the one this site did not answer

A sentence is enough. It is answered by the people who built the engine, and if the honest answer is a demonstration, we will offer one on your own data.