LmCast :: Stay tuned in

AI models need more data about biology, and OpenAI is paying to create it

Recorded: Sept. 15, 2026, 2:48 p.m.

Original Summarized

AI models need more data about biology, and OpenAI is paying to create it | MIT Technology Review

You need to enable JavaScript to view this site.

Skip to ContentMenuMIT Technology ReviewMIT Technology ReviewThe Big StoryArtificial intelligenceBiotech & healthClimate & energy10 Breakthrough TechnologiesEmTech Future live eventThe Kids issueMenuMIT Technology ReviewMIT Technology ReviewThe Big StoryArtificial intelligenceBiotech & healthClimate & energy10 Breakthrough TechnologiesEmTech Future live eventThe Kids issueBiotechnology and healthAI models need more data about biology, and OpenAI is paying to create itThe OpenAI Foundation is funding a new effort called Data for Public Health.
By Antonio Regaladoarchive pageSeptember 15, 2026Stephanie Arnett/MIT Technology Review | Adobe Stock Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies. By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She called these documents “biotech’s lost archive” and said they could be used to help train AIs that would act as powerful copilots in the often opaque drug approval process.  Today the OpenAI Foundation, the nonprofit parent of OpenAI, said it would fund her idea as part of a new effort it calls Public Data for Health, which aims to help artificial intelligence make big leaps in medicine by paying to create “high-quality scientific datasets.”
The basic idea is that AI isn’t going to be capable of making important breakthroughs in curing disease unless researchers can feed the models much more information than they have so far.  “Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology,” says Morgan Levine, a former vice president for computation at Altos Labs, a longevity company.
Related StoryColossal Biosciences is growing chickens in a 3D-printed artificial eggshellRead next In its initial round of data grants, the OpenAI Foundation also announced that it would give $40 million to a program to collect data about novel cancer vaccines at the University of North Carolina, Chapel Hill, and support OpenAdmet, a group that runs competitions in which researchers try to predict drug effects.  Teslo’s idea for a biotech archive received $500,000 and will be pursued by 1Day Sooner, an advocacy group representing clinical trial volunteers, which she advises. “We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data,” the OpenAI Foundation said in a statement. OpenAI started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion. Because the foundation holds a 26% equity stake in OpenAI, it is now be on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. (By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.)   Making good use of that kind of money will not be easy. The foundation, based in San Francisco, is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition, an organization that helps patients get access to drugs for hepatitis C. OpenAI’s charitable efforts come even as apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon. Those fears have been stoked by AI company insiders, some of whom say the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to “slow the pace at which we improve the capabilities of AI models” so that risk prevention can catch up.

Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.” “We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year.  The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each. Related StoryHere’s why Elon Musk lost his suit against OpenAIRead next His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts.  Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted.  Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed.  The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug. According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market. “People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.”  by Antonio RegaladoShareShare story on linkedinShare story on facebookShare story on emailPopularA fundamental flaw leaves LLMs strikingly vulnerable to attackWill Douglas HeavenAI is more likely than humans to form biases when hiringMichelle KimHere’s why AI agents lie and cheat to reach their goalsGrace HuckinsAI’s recursive self-improvement might not come so quickly after allMichelle KimDeep DiveBiotechnology and healthA startup claims it’s found a drug to make your blood youngGeneration Lab claims its drug combo can “stop the spread of aging” around the body. And it’s looking for influencers to give it a try.
By Antonio Regaladoarchive pageMontana’s plan to become an experimental medical hub just pushed forwardThe state’s effort to expand the “right to try” is making headway, and the first drugs are about to be reviewed.
By Jessica Hamzelouarchive pageThere’s a lot of hype around perimenopause. Don’t buy it.Discussions of the life stage are often clouded by misinformation.
By Jessica Hamzelouarchive pageSupercooled kidneys have been transplanted into pigs in a “landmark achievement”Kidneys kept at subzero temperatures in pressure-controlled containers can be stored for days before transplantation, raising hopes for longer-term storage of donated human organs.
By Jessica Hamzelouarchive pageStay connectedIllustration by Rose WongGet the latest updates fromMIT Technology ReviewDiscover special offers, top stories,
upcoming events, and more.Enter your emailPrivacy PolicyThank you for submitting your email!Explore more newslettersIt looks like something went wrong.
We’re having trouble saving your preferences.
Try refreshing this page and updating them one
more time. If you continue to get this message,
reach out to us at
customer-service@technologyreview.com with a list of newsletters you’d like to receive.The latest iteration of a legacyFounded at the Massachusetts Institute of Technology in 1899, MIT Technology Review is a world-renowned, independent media company whose insight, analysis, reviews, interviews and live events explain the newest technologies and their commercial, social and political impact.READ ABOUT OUR HISTORYAdvertise with MIT Technology ReviewElevate your brand to the forefront of conversation around emerging technologies that are radically transforming business. From event sponsorships to custom content to visually arresting video storytelling, advertising with MIT Technology Review creates opportunities for your brand to resonate with an unmatched audience of technology and business elite.ADVERTISE WITH US© 2026 MIT Technology ReviewAboutAbout usCareersCustom contentAdvertise with usInternational EditionsRepublishingMIT Alumni NewsHelpHelp & FAQMy subscriptionEditorial guidelinesPrivacy policyTerms of ServiceWrite for usContact uslinkedin opens in a new windowinstagram opens in a new windowreddit opens in a new windowfacebook opens in a new windowrss opens in a new window

Artificial intelligence models require significantly more biological data to achieve major breakthroughs in medicine, a necessity that is driving funding initiatives from organizations like the OpenAI Foundation. This push stems from the recognition that data represents the primary bottleneck in successfully applying artificial intelligence to biological sciences. Ruxandra Teslo proposed utilizing data from failed biotechnology companies, which she termed "biotech’s lost archive," arguing that detailed regulatory filings, manufacturing strategies, and safety data could be leveraged to train AI systems that function as powerful copilots in the often opaque drug approval process.

The OpenAI Foundation has responded to this vision by launching an initiative called Data for Public Health, aimed at funding the creation of "high-quality scientific datasets" to facilitate these advancements. This effort is based on the principle that combining the intelligence of new AI models with increased real-world observations, or more data, is essential for achieving breakthroughs in curing diseases. Morgan Levine, a former vice president for computation at Altos Labs, affirmed that data is the most significant obstacle to successfully applying AI to biology.

Beyond this initiative, the Foundation has allocated substantial funds to related research, including $40 million to collect data concerning novel cancer vaccines at the University of North Carolina, Chapel Hill, and support competitions such as OpenAdmet, where researchers attempt to predict drug effects. Furthermore, the concept of accessing proprietary data is being explored; the grant of $500,000 provided to 1Day Sooner, an advocacy group for clinical trial volunteers, aims to demonstrate the feasibility of obtaining datasets from bankrupt companies, suggesting that nonexclusive copies of company data might be obtainable for a relatively low cost.

The specific type of data sought, known as common technical documents, includes the extensive back-and-forth between companies and regulators, alongside detailed scientific and medical measurements, offering a comprehensive view of what is known about a drug. Teslo posits that such a stockpile of information could transform an AI into a regulatory expert, potentially accelerating cures to market, particularly because the clinical development process is largely a black box, despite accounting for most of the time and money spent in clinical trials. This reliance on verifiable, comprehensive data is critical because the journey from drug innovation to market approval is complex, and the messy reality of the regulatory process needs to be integrated with AI capabilities.

The context surrounding this data focus is increasingly intertwined with broader concerns regarding the development of artificial intelligence. While the focus is on medical applications, the evolution of AI models operates alongside anxieties about the potential risks associated with runaway AI and the imperative to slow down capability improvements for risk prevention. Leaders such as Sam Altman and Elon Musk have endorsed calls for moderating the pace of AI advancement to ensure safety. The OpenAI Foundation operates with a mission to ensure artificial intelligence benefits all of humanity, working through external nonprofits and research institutions to achieve this goal, managing significant charitable resources while navigating these complex technological and ethical challenges.