The DeWitt clause

Why database licences spent forty years forbidding you to publish a benchmark, and what that did to the evidence.

8 min read

There is a reason you have read a hundred vendor benchmarks and almost no independent ones, and it is older and more deliberate than «benchmarks are expensive». For most of the history of this industry, publishing one was a breach of contract.

1982

David DeWitt, then at the University of Wisconsin, published a benchmark study comparing database systems. Oracle’s showed poorly. Larry Ellison called the department chair and demanded DeWitt be fired. Wisconsin refused.

What happened next is the part that mattered. Ellison banned Oracle from hiring University of Wisconsin students, and Oracle changed its licence so that nobody could publish a benchmark naming Oracle without permission. Nearly every other database vendor adopted the same term, and the anti-benchmarking clause became a de facto industry standard — known ever since by the name of the academic it was written to punish.

The clauses spread beyond databases, into compilers, and later into cloud services.

What the clause actually says

The form is consistent. Microsoft’s SQL Server licence has said you may not disclose the results of any benchmark test without Microsoft’s prior written approval.

Note what that does and does not forbid. It does not forbid measuring — you may benchmark all you like in private. It forbids publishing. The prohibition is aimed precisely and only at the part where other people find out.

And note who it binds. Not the vendor. The vendor may publish whatever numbers it likes about itself, and about competitors whose licences it never accepted. The clause creates an asymmetry in which the only party permitted to publish a comparison is the party with the most to gain from its outcome.

That asymmetry is the answer to a question I had before reading any of this: why is the benchmark literature in this field so overwhelmingly vendor-authored? Partly cost. But partly because for decades the alternative was legally actionable.

The thaw, and its limits

This is no longer the whole picture, and the change is recent enough to be worth dating.

In November 2021 — the same month as the Databricks/Snowflake exchange, which is not a coincidence — Databricks removed the DeWitt clause from its licence and published an argument for the industry doing likewise. SingleStore did the same. The case they made was that the clause suppresses exactly the scrutiny that would make performance claims meaningful.

Two qualifications keep this from being a happy ending. The first is that plenty of licences still carry it, and the burden falls on whoever wants to publish to read the terms of every system they name. The second is subtler: cloud providers have adopted reciprocal versions, permitting benchmarking on condition that competitors permit it too. That is better than a flat prohibition and it is not the same as freedom to publish — it makes your right to report a measurement contingent on somebody else’s licence terms.

Why this sits on this site

Three reasons, and the third is the real one.

It explains the shape of the evidence. When you notice that almost every comparison you can find was written by a party to it, the natural conclusion is that independent writers are lazy or broke. Some of that is true. But a legal regime that made independent publication a breach of contract for forty years is a better explanation for a forty-year gap.

It sets an expectation for what I can cover. Systems whose licences still forbid publication are systems this site will either not name or will approach carefully, and I would rather you know that the absence of a particular product from a comparison here may be a legal fact rather than a technical judgement.

And it is a reminder that «nobody has measured this» is a claim about the world, not about the question. When I write that a number is unpublished — commit latency across Iceberg catalogs, the idle cost of an orchestrator, the point where recursive CTEs give up — the reason is usually that measuring it is unglamorous work nobody funded. Sometimes the reason is that somebody made sure it would not be published. Both are gaps. Only one of them is an accident.