

The US government, Google, and Meta are joining forces with Zuckerberg-backed nonprofit Biohub in a USD 1.8 billion effort to build large-scale biological datasets for artificial intelligence research. The initiative, announced on Wednesday (October 7, 2026), aims to generate enough data on how cells behave under different conditions to help scientists develop predictive AI models and potentially shorten the drug development process.
The project, called the Virtual Biology Initiative, brings together Meta, Google DeepMind, drug-discovery company Isomorphic Labs, the US Department of Energy and the National Institutes of Health. Biohub was founded with backing from Meta CEO Mark Zuckerberg and his wife, Priscilla Chan.
The latest commitments include more than USD 500 million from the Department of Energy over five years for laboratory measurements, modeling, and computing. Meta, Google DeepMind and Isomorphic Labs will jointly contribute USD 300 million.
The NIH will also coordinate and standardize datasets and repositories created through more than USD 500 million in previous federal funding so they can be used for AI training. Biohub had committed another USD 500 million to the initiative earlier this year.
The combined effort aims to address a major limitation in applying AI to biology: far less high-quality, structured biological data is available than the enormous datasets used to train many modern AI systems.
Biohub plans to collect information showing how cells respond to different environmental and biological changes. Researchers will use techniques including spatial transcriptomics, which maps molecular activity within intact tissue, along with large-scale experiments that measure cellular responses.
The long-term goal is to create predictive models that can simulate aspects of cellular behavior. Scientists could potentially use those models to test ideas computationally before committing time and resources to laboratory experiments.
Biohub's head of science Alex Rives said existing datasets contain hundreds of millions of cells, while accurate predictive models may eventually require billions or even trillions of observations.
Also Read: How Mark Zuckerberg is Building the Future of AI-Powered Social Media
The initiative is being presented as an open-science effort, but commercial partners will receive temporary early access to datasets they help fund. After these embargo periods, the data will be released as a public scientific resource. Government-funded work will not have the same restrictions.
Biohub expects the first dataset to be ready in about a year and aims to develop increasingly accurate predictive models within five years. The initiative comes as other AI companies and research organizations are also investing heavily in biological datasets and drug discovery. Anthropic has expanded its biology work, while the OpenAI Foundation has launched a grant program to support biological and medical datasets.