AI models need more data about biology, and OpenAI is paying to create it
Recorded: Sept. 15, 2026, 2:48 p.m.
| Original | Summarized |
AI models need more data about biology, and OpenAI is paying to create it | MIT Technology Review You need to enable JavaScript to view this site. Skip to ContentMenuMIT Technology ReviewMIT Technology ReviewThe Big StoryArtificial intelligenceBiotech & healthClimate & energy10 Breakthrough TechnologiesEmTech Future live eventThe Kids issueMenuMIT Technology ReviewMIT Technology ReviewThe Big StoryArtificial intelligenceBiotech & healthClimate & energy10 Breakthrough TechnologiesEmTech Future live eventThe Kids issueBiotechnology and healthAI models need more data about biology, and OpenAI is paying to create itThe OpenAI Foundation is funding a new effort called Data for Public Health. Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.” “We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year. The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each. Related StoryHere’s why Elon Musk lost his suit against OpenAIRead next His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts. Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted. Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed. The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug. According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market. “People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.” by Antonio RegaladoShareShare story on linkedinShare story on facebookShare story on emailPopularA fundamental flaw leaves LLMs strikingly vulnerable to attackWill Douglas HeavenAI is more likely than humans to form biases when hiringMichelle KimHere’s why AI agents lie and cheat to reach their goalsGrace HuckinsAI’s recursive self-improvement might not come so quickly after allMichelle KimDeep DiveBiotechnology and healthA startup claims it’s found a drug to make your blood youngGeneration Lab claims its drug combo can “stop the spread of aging” around the body. And it’s looking for influencers to give it a try. |
Artificial intelligence models require significantly more biological data to achieve major breakthroughs in medicine, a necessity that is driving funding initiatives from organizations like the OpenAI Foundation. This push stems from the recognition that data represents the primary bottleneck in successfully applying artificial intelligence to biological sciences. Ruxandra Teslo proposed utilizing data from failed biotechnology companies, which she termed "biotech’s lost archive," arguing that detailed regulatory filings, manufacturing strategies, and safety data could be leveraged to train AI systems that function as powerful copilots in the often opaque drug approval process. The OpenAI Foundation has responded to this vision by launching an initiative called Data for Public Health, aimed at funding the creation of "high-quality scientific datasets" to facilitate these advancements. This effort is based on the principle that combining the intelligence of new AI models with increased real-world observations, or more data, is essential for achieving breakthroughs in curing diseases. Morgan Levine, a former vice president for computation at Altos Labs, affirmed that data is the most significant obstacle to successfully applying AI to biology. Beyond this initiative, the Foundation has allocated substantial funds to related research, including $40 million to collect data concerning novel cancer vaccines at the University of North Carolina, Chapel Hill, and support competitions such as OpenAdmet, where researchers attempt to predict drug effects. Furthermore, the concept of accessing proprietary data is being explored; the grant of $500,000 provided to 1Day Sooner, an advocacy group for clinical trial volunteers, aims to demonstrate the feasibility of obtaining datasets from bankrupt companies, suggesting that nonexclusive copies of company data might be obtainable for a relatively low cost. The specific type of data sought, known as common technical documents, includes the extensive back-and-forth between companies and regulators, alongside detailed scientific and medical measurements, offering a comprehensive view of what is known about a drug. Teslo posits that such a stockpile of information could transform an AI into a regulatory expert, potentially accelerating cures to market, particularly because the clinical development process is largely a black box, despite accounting for most of the time and money spent in clinical trials. This reliance on verifiable, comprehensive data is critical because the journey from drug innovation to market approval is complex, and the messy reality of the regulatory process needs to be integrated with AI capabilities. The context surrounding this data focus is increasingly intertwined with broader concerns regarding the development of artificial intelligence. While the focus is on medical applications, the evolution of AI models operates alongside anxieties about the potential risks associated with runaway AI and the imperative to slow down capability improvements for risk prevention. Leaders such as Sam Altman and Elon Musk have endorsed calls for moderating the pace of AI advancement to ensure safety. The OpenAI Foundation operates with a mission to ensure artificial intelligence benefits all of humanity, working through external nonprofits and research institutions to achieve this goal, managing significant charitable resources while navigating these complex technological and ethical challenges. |