Earlier this year I argued that no experiment is an island: that the value of biological data shows up in context, not in isolation. The most common response I heard was some version of: yes, obviously, so why is it still so hard? That's exactly the right question, and I want to answer it head-on, because it points straight at the biggest shift I see coming to our field.
For two decades, the life sciences have poured extraordinary effort into storing data. We built repositories, standards, and pipelines to capture ever-larger datasets. That work matters enormously. But somewhere along the way we mistook storage for progress. We now have vast amounts of biology sitting in archives that are, functionally, write-only: deposited, cited once, and never truly reasoned across again.
Storage answers the question "where is the data?" It was never designed to answer the questions scientists actually have: what does all of this, taken together, mean? The gap between those two verbs, storing versus reasoning, is where I believe the next decade of value in our industry is hiding.
I've started using a phrase inside our walls that I want to say out loud: data intelligence. Not as a product name, but as a direction. By it I mean the ability to treat a large, curated body of biological data as something you can genuinely interrogate: ask a real question and get an answer that draws across thousands of experiments, with the provenance to back it up. A living, queryable body of knowledge, rather than a graveyard of files.
Three things have to come together for that to be real. First, the data has to be harmonized and trustworthy, the foundation we've been laying for years. Second, it has to be connected: related experiments, entities, and findings linked so you can traverse them, the way a knowledge graph lets you follow biology from a gene to a pathway to a disease. Third, and increasingly, it has to be legible to AI, because the scale is far past what any human can hold in their head, and the tools to help have finally grown genuinely capable.
I want to be careful here, because "AI across all your data" is a promise our industry has heard before and rightly learned to distrust. The difference between hype and help is provenance. An answer you can't trace is a liability, not an asset, especially in a field where the decisions carry real consequences. So the kind of data intelligence I believe in is inseparable from attestation: every answer able to show its work, back to the underlying evidence and to how that evidence was produced. Intelligence you can audit. That's the only version worth building.
This is a genuine pivot, and I don't want to undersell it. It reframes what a platform like ours is for: not a place you run an analysis and leave, but a place where your data joins a larger, growing whole and becomes more useful for having done so. You'll see us keep pulling more data into one rigorous, interoperable home, and keep investing in the connections and the intelligence layered on top of it.
From storage to intelligence. That's the road. It runs directly through everything we've built so far, and it's exactly where we're pointed next. As always, thank you for building it with us.