Biohub gathers $1.8B from Meta, Google DeepMind and the US government to build open AI training data for biology
Meta, Google DeepMind and Isomorphic Labs are putting in a combined $300M. The Department of Energy has pledged more than $500M over five years. The NIH will coordinate data work that builds on about $500M of earlier federal funding. Nvidia, Arc Institute and Tahoe Therapeutics are among the partners. The stated goal is a predictive 'virtual cell' model within five years. Biohub is backed by Mark Zuckerberg and Priscilla Chan. The program is meant to produce open, standardized biological datasets built for training models, not for single studies. It extends the virtual cell work already under way at Arc Institute, and it fits a wider push by labs such as Isomorphic to use foundation models for drug discovery. For founders in AI biology, the scarce input has been data, not model architecture. Large open cell datasets paid for by government and big tech would make it cheaper for startups to train bio foundation models. They would also weaken the proprietary data moats some companies are building now.