#C2410C
hex37

Plain-English explanations of the stories the internet is arguing about.

AI · 5 min

Most of AI's Billion-Dollar Startups Publish Almost No Research

A study of the AI companies valued above $1 billion found that only about half have contributed anything to the scientific literature. The famous labs turn out not to be the quiet ones — and the reasons the others stay silent say a lot about how the industry now works.

The library of the University of the Basque Country (Vitoria-Gasteiz) receives an impressive selection of scientific journals.
Photo: Vmenkov / CC BY-SA 3.0 · source

Somewhere between the laboratory and the product launch, a large part of the artificial intelligence industry stopped writing things down.

That is the finding at the centre of a report in Science on the scientific output of AI's "unicorns" — the private companies whose investors value them at more than a billion dollars. Researchers counted how often these firms appear in the published scientific record: papers submitted to journals and conferences, reviewed by other specialists in the field, and then made available for anyone to read, check and build on. Roughly half of the companies barely register at all.

That number can be read two ways, and both readings are honest. Half of a cohort of fast-growing commercial companies contributing anything to public science is arguably more than you would expect from any other industry. It is also a striking figure for a field whose entire existence rests on work that earlier researchers gave away.

The famous labs are not the quiet ones

The instinct on reading a headline like this is to assume it means OpenAI and Anthropic have gone dark. The measurement says close to the opposite. Both companies appear among the publishers, as does Hugging Face, the company that runs the main public repository where AI models and training datasets are shared. On the study's ranking of cumulative citations — how often other researchers refer back to a company's papers — OpenAI, the maker of ChatGPT, sits at the top. The rest of the upper end is a mix of Chinese computer-vision firms, self-driving-car companies and outfits applying machine learning to drug discovery.

Two caveats matter here. Citations are a proxy for influence, not for truth: a paper can accumulate references because it is important, because it is convenient to cite, or because it is wrong in an interesting way. And the study covers startups only. Google, Microsoft and Meta are excluded by definition, despite employing some of the largest research operations in the field. So this is not a measurement of whether AI as a whole publishes. It is a measurement of what happens to a research culture when it is absorbed into venture-funded companies.

Publishing hands rivals a six-month head start

The commercial logic against publishing is not subtle. A team that spends six months finding a technique that makes a model faster or cheaper can describe it in a paper, at which point a better-resourced competitor can re-implement it in a fraction of that time. Nothing comes back the other way. Unlike an industry built on physical products, where a rival has to buy your device and take it apart, AI research is transferable the moment it is legible. Description is the whole product.

The publishing system itself does not help. Getting a result into a prestigious journal or conference can take a year or more of review and revision — an eternity in a field where the state of the art moves quarterly. Many researchers now skip that queue entirely and post preprints: papers put online publicly before any formal review. That is faster, but it shifts the burden of quality control onto readers who mostly do not have the compute budget to check anything.

What often results is a third category — the marketing paper. Enough technical detail to impress investors and recruit engineers, not enough for anyone to reproduce the result. The paper arrives attached to the fundraising deck rather than to the field.

When the paper is a blog post

The gap has been filled by company blog posts announcing breakthroughs, complete with charts and benchmark scores, and with no methods section, no disclosed experimental setup and no independent verification. Benchmarks — standardised tests used to compare AI systems — can be chosen after the results are in. Claims spread by the dynamics of social media rather than of scholarship: the striking framing travels, the caveat does not.

This has a peculiar second-order effect. The blog posts, papers and confident summaries generated in this environment become part of the text used to train the next generation of models, which in turn help produce the next round of write-ups. Unverified claims do not simply sit in an archive; they get absorbed.

It is worth resisting the temptation to romanticise the alternative. Academic publishing has its own well-documented pathologies: enormous volumes of low-quality output produced to satisfy institutional metrics, review processes that function more as credentialing than communication, and commercial publishers extracting fees from work funded by the public. Peer review is a filter, not a guarantee. The problem is not that the old system worked perfectly. It is that the new one has no filter at all.

Trade secrets, unlike patents, never expire

The deeper shift is legal rather than cultural. Patents were designed as a bargain: an inventor discloses how something works, in full, and in exchange gets a time-limited monopoly on using it. Society gets the knowledge; the inventor gets the years of exclusivity. That bargain assumed the invention would eventually have to be revealed in order to be sold.

When the product is a service delivered through an interface — a chatbot, an API, a model you rent by the token — nothing needs to be revealed. There is no device to buy and dismantle. Every improvement can remain a trade secret indefinitely, and trade secrets have no expiry date.

The knock-on effects reach past the companies. Ideas cross-pollinate more slowly when each firm's advances are invisible to the others, which means the field as a whole compounds more slowly. And the people doing the work end up holding skills that are specific to one employer's undisclosed systems, which makes those skills harder to carry elsewhere and weakens their position when negotiating.

None of this makes the startups unreasonable. Most AI companies founded in the past three years are building products on models someone else trained; expecting scientific papers from them is a category error, like asking a restaurant to publish agronomy research. Productising other people's discoveries is a legitimate job.

But the industry is now large enough, and well-funded enough, that the question is no longer whether individual founders are behaving sensibly. It is whether a field that grew by reading everyone else's homework can keep growing once it stops handing in its own.

Questions

Does this mean OpenAI and Anthropic have stopped publishing research?

No. Both appear among the companies that do publish, and OpenAI ranks at the top of the study's citation counts. The finding concerns the wider population of billion-dollar AI startups, many of which are commercial product companies rather than research labs.

Why does it matter whether a private company publishes its research?

Published methods let other researchers verify claims, build on them and avoid repeating dead ends. When advances stay secret, the field depends on duplicated effort, and workers' skills become tied to one employer's undisclosed systems.

Isn't a company blog post about a new model the same thing as a paper?

Not really. A peer-reviewed paper is checked by independent specialists and is meant to contain enough detail for others to reproduce the result. A blog post typically reports benchmark scores without disclosing the experimental setup, so the claims cannot be independently confirmed.

Why weren't Google and Meta included in the study?

The study looked specifically at startups valued above a billion dollars. The large established technology companies fall outside that definition, even though they run some of the field's biggest research operations.

Read the original at science.org →