AI Can Out-Invent Big Pharma. It Still Can't Out-Build It.
- Jul 10
- 6 min read
Updated: Jul 13
In drug discovery, a good algorithm can beat a bigger budget. In drug manufacturing, AI doesn't even get a seat at the table until someone else has already solved the capital problem.
Prepared by Richstorm.co

Key Takeaways
AI's usefulness in pharma depends entirely on the kind of problem being solved, not on how advanced or well-funded a company is.
In drug discovery, AI directly attacks the problem of searching for the right molecule, which is why small, well-run teams can already compete with giants there, no factory required.
Rentosertib, an AI-designed drug from Insilico Medicine, posted positive Phase IIa trial data published in Nature Medicine in June 2025, real proof that this kind of discovery works.
In drug manufacturing, AI can't attack the problem directly at all, because it needs years of proprietary production data that only exists once a company has already built and run an expensive facility.
Companies are racing to build that manufacturing capacity for ordinary business reasons that have nothing to do with AI; once they own it, AI becomes a compounding advantage only they get to use.
One Technology, Two Very Different Jobs
AI in pharma gets talked about as one story. It's really one technology doing two completely different jobs, and it only succeeds at one of them on its own merits. The real reason comes down to where the data lives.
Drug discovery runs on data that's largely public. A patent legally requires disclosing a drug's chemical structure in exchange for exclusivity, and decades of published papers and open databases like PubChem and the Protein Data Bank have put millions of molecular structures and activity results within reach of anyone. That's exactly why two university research teams, at Huazhong University of Science and Technology and South China University of Technology, could each build a working ADC prediction model with no pharmaceutical manufacturing operation of their own: they trained on published data, not on any single company's private lab results.
Manufacturing has no real equivalent. A drug's chemical structure has to be disclosed to regulators to get approved, and once approved it can often be worked out just by analyzing the product itself, so patenting it is really the only way to protect it. A manufacturing process is different: it's frequently very hard to reverse-engineer from the finished drug, so companies usually choose to protect it as a trade secret instead of patenting it, since a trade secret can last indefinitely while a patent expires in 20 years and requires disclosing exactly how the process works.
FDA does receive detailed manufacturing and batch data as part of every drug approval, but its own regulations classify that information as trade secret or confidential commercial information and exempt it from public disclosure by name, even under a Freedom of Information Act request. There's no public database of bioreactor conditions or batch yield data the way there is for molecular structures, because that data was never intended to become public in the first place. AI is only as good as the data it can reach, and manufacturing data simply doesn't exist anywhere outside the walls of whichever company generated it by running its own facility.
Where AI Wins on Its Own: Drug Discovery
Small Molecules: The Most Mature AI Track Record
Insilico Medicine's rentosertib, an oral small-molecule drug targeting a protein involved in lung scarring, went from initial target identification to a preclinical candidate in just 18 months using the company's PandaOmics and Chemistry42 platforms. In a 71-patient Phase IIa trial published in Nature Medicine in June 2025, patients on the drug's 60 mg daily dose showed a mean lung function improvement of 98.4 mL, compared to a 20.3 mL decline in the group given a placebo. It is the first drug with a fully AI-designed target and compound to receive an official drug name from the United States Adopted Names system, granted in March 2025, a real regulatory milestone, not just a research claim.
The same public-data pattern shows up here too: the target was identified by mining public patents, papers, and clinical trial databases, and the molecule itself was designed using protein structures from AlphaFold, which was trained on the public Protein Data Bank. Neither step required Insilico's own manufacturing data.
Antibodies and Bispecifics: Real Clinical Momentum, Not Yet Full Proof
The biotech company LabGenius has used a fully AI-driven design loop to develop highly selective bispecific T-cell engager antibodies and expects to file for permission to begin human trials in 2026. A language-model-based design approach developed at MIT achieved sub-nanomolar binding strength in 99% of the antibody candidates in its top-performing design batch, and during the COVID-19 pandemic, Just-Evotec Biologics used an AI-based platform to identify functional therapeutic antibodies active against multiple variants of the virus.
RNA and mRNA Therapeutics: Strong in the Lab, Unproven in Patients
Researchers at Stanford University trained a model that correctly classified which lipid nanoparticle formulations would work well 98% of the time on unseen data. Separately, a team at the University of Macau used AI to screen nearly 20 million candidate lipid molecules, and in mouse testing, several of the newly designed candidates matched or outperformed an already-proven delivery formulation used in existing approved therapies.
Antibody-Drug Conjugates: Good at Predicting, Not Yet a Proven Drug
Researchers at Huazhong University of Science and Technology built a model called DumplingGNN that correctly identified active versus inactive ADC payload molecules with 91.5% accuracy. A separate team at South China University of Technology built ADCNet, achieving roughly 87% accuracy, and made it available as a free public tool. What's still missing is a public case of either tool being credited with picking a payload that became a successful, approved drug, but the underlying predictive science is real.
Cell Therapy and Gene Therapy: A Different Kind of Problem Already
These two modalities are the hinge point of this piece. The target is largely defined by the disease itself before development starts, a known cancer antigen for CAR-T, a patient's own diagnosed faulty gene for gene therapy, so there's little of the large-scale candidate searching that makes AI so useful elsewhere. The real bottleneck here has always been manufacturing, not discovery, which is exactly where the story changes.
Where AI Doesn't Compete at All: Drug Manufacturing
It helps to see how this plays out in a specific modality rather than just in the abstract. Antibody-drug conjugates make a good example. The reason companies are racing to expand ADC manufacturing capacity has nothing to do with AI: ADCs are already a proven, fast-growing business on their own, with Enhertu alone generating close to $5 billion a year, the ADC pipeline growing over 25% annually since 2020, and specialized manufacturing capacity chronically short of demand for years. That commercial case would exist with or without AI.
The concentration this creates is real. Pfizer, Lonza, and Catalent together run seven of the twelve global facilities capable of conjugating a toxic payload onto an antibody at commercial scale, and Lonza alone already manufactures more than half of all ADCs approved to date. Every major player is racing to expand further, including a $10.5 billion Pfizer commitment in May 2026 to co-develop a 12-drug ADC portfolio with Innovent Biologics, purely to capture more of that existing demand.
AI only enters the picture once that capacity already exists. Kriya Therapeutics used AI to screen roughly 80 viral vector construct designs for manufacturability, but only after spending over a billion dollars building the facility that generated the data to do it. Lonza applies predictive modeling to select the right process, cell clone, and molecule, and Samsung Biologics' newest plant uses AI and digital twin technology to forecast batch yield. Samsung Biologics has said outright that every major CDMO competitor now runs the same playbook: build capacity for commercial reasons first, then layer AI on top of the data that capacity generates.
What This Means for Where to Look
These are two different questions, and they deserve different scrutiny. A small or mid-size biotech's AI story is genuinely credible when it's about designing molecules, antibodies, or delivery systems, since that race is winnable on science alone, without owning a factory first. A company's claimed AI manufacturing advantage is a different kind of claim entirely: it's really a capital story wearing an AI label, and it's only real if that company already won the infrastructure race for other reasons first. Asking which one you're looking at is the difference between evaluating a company's science and evaluating its balance sheet.
A companion piece on RichStorm ranks every major drug modality, from small molecules to gene therapy, by how difficult it actually is to build, and explains why the pharmaceutical industry needs so many different approaches in the first place.

