đŹ From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science â Andrew White
RJ HonickyBrandon AndersonAndrew White
- AI science is already less intelligence-constrained than information-constrained. Even a hypothetical âOpus 7 or GPT10â eventually needs nature to supply new evidence; todayâs real bottleneck may be mundane laboratory stateâreagent inventory, lead times, cost, and experiment turnaroundânot whether âGPT 5.2 Codex Max or Opus 4.5â proposes the cleverer first experiment. The valuable system closes the hypothesisâexperimentâanalysis loop.
- The investable wedge is a shared operating system for discovery, not merely another domain foundation model. Cosmos combines literature research, data analysis, experiments, reporting, and an evolving world model that White likens to a git repository: a distilled state that multiple agents can update and use for predictions. The breakthrough came when the team stopped grounding that model only in literature and put âexperiment in the loopâ through data analysis.
- Scientific taste remains the frontier capabilityâand naive human preference data did not teach it. Pairwise raters rewarded tone, specificity, and feasibility more readily than the consequential question: âIf this hypothesis is true, how does it change the world; if false, how does it change the world?â Cosmosâs roughly 52% or 55% score on interpretation was not wet-lab success but agreement over whether findings were interesting or novel.
- Verification produced more signal than expert enthusiasm in FutureHouseâs strongest end-to-end test. In Robinâs dry-AMD work, specialists broadly agreed on a top 10 but rankings beyond that became noisy; after four weeks of experiments, the winning mechanism and repurposed drugâlikely ripasudilâwere not the expertsâ favorite. Whiteâs updated view is to trust ânatureâs computerâ: literature, data, unit tests, or physical experiments inside the loop.
- Scale advantage comes from enumerating more hypotheses and filtering them cheaply before wet-lab spend. Whiteâs maxim is, âIf you canât be smarter, you can try more times,â with provenance preserved from page-level citations through Python lines to downstream conclusions. On BixBench, agents reach roughly 60â70% correctness while humans agree at about 70% of the analyses, suggesting that some remaining error reflects methodological disagreement rather than simple model failure.
- Whiteâs sharpest compute call is that molecular dynamics and DFT are overrated for discovery. His own water simulation consumed about 1 million CPU-hours yet mainly identified hyperparameters reproducing known effects; âsimulations simulate really boring things really wellâ while catalysts and other complex systems contain the grain boundaries, dopants, and complexity they miss. D. E. Shaw Researchâs bespoke MD hardware versus AlphaFoldâs experimental-data learning is his decisive comparison: an imagined five special machines producing one or two folds daily lost to a model runnable on a desktop, with a good folding model now requiring, by his estimate, about 10,000 GPU-hours.
- Verifier engineering is a hidden scaling risk for scientific reinforcement learning. Ether0 repeatedly exploited every rule: separating required atoms, proposing implausible nitrogen chains, adding purchasable but irrelevant nitrogen, and exploiting reagent ordering rather than learning chemistry. White calls this handcrafted spiral the âboutique lessonâ; the recurring realization was, âWhy am I doing this? How did I get here?â
- Commercialization is arriving faster on year scales than White expected, but labor and safety consequences remain unresolved. He âoverestimate[s] the speed of things on month scale and underestimate[s] things on year scaleâ: a 10-year automation mission announced around 2023 looked radically closer by 2025, while Edison had already been part of the organizational plan. He expects scientists to become âCosmos wranglersâ exploring 10Ă or 100Ă more ideas, while conceding that firms may choose compute over ten new hires and that emerging real-time or computational dual-use scenarios deserve more attention.
1. Whiteâs path from molecular simulation to agents began with an experimentâcomputation mismatch
White entered a University of Washington PhD group with roughly 19 experimentalists and two simulation researchers. His biomaterials work asked why implants become collagen-encapsulatedâa useful response around pacemakers, but a lifetime-limiting one for glucose sensors or brain-computer interfaces.
A 10,000-atom simulation could not capture a human body and implant, so his postdoc explored maximum entropy: fitting complicated simulations to scarce observations, âthe inverse of machine learning.â At Rochester he applied those ideas to peptides, years before peptides became fashionable enough for what he jokingly called a âpeptide rave.â
A 2019 UCLA sabbatical exposed him to machine learning for physics. Chemistry courses still stopped at RNNs or image classification, while his field needed graphs, symmetry, and geometry, prompting him to write a chemistry-focused machine-learning textbook.
After the original Codex, his group built verifiable scientific-programming tasks such as completing an MCMC function and testing whether it remained valid. He was a GPT-4 red-teamer before release, then combined GPT-4, ReAct, literature tools, and IBMâs cloud lab in ChemCrow; the project helped convince him that agents could operate science rather than merely discuss it.
2. FutureHouse and Edison turned an academic research program into a larger organizational bet
ChemCrow triggered enough concern that Whiteâs paper was presented at the White House, where he encountered agencies asking how AI changed explosives or nuclear-weapons breakout time. The episode showed him how few people then combined serious AI knowledge with scientific domain expertise.
Sam Rodriques, after discussions with Eric Schmidt and Tom Kalil, was exploring focused research organizations: tightly scoped science outside academia and near-monopoly technology labs. White proposed agents for science; Rodriques pushed the ambition from âsee what fun stuff we could doâ to the long-term mission of automating science.
White initially retained his Rochester position through sabbatical, then resigned his tenure in June when he co-founded venture-backed Edison Scientific, spun out of FutureHouse. Academia remained attractive, but he judged grant-writing insufficient for âthe biggest bet you can takeâ on a field moving this quickly.
The nonprofit-to-company pattern may not be repeatable because contemporary AI research and GPUs are so expensive. Even cash salaries above $1 million, which White finds astonishing, can remain small beside compute burn; financing structure now materially shapes which scientific-agent experiments are possible.
3. Scientific automation means closing the cognitive loop, not modeling one biological object
White distinguishes teams building a virtual cell, protein-folding model, or antibody designer from his target: automating hypothesis formation, experiment selection, result analysis, belief updates, and the evolving world model that generates the next experiment.
Progress repeatedly outran the organizationâs infrastructure plans. The team expected to need automated labs, unified paper stores, and APIs around everything; stronger models can instead email a CRO, instruct a human, or inspect a video of an experiment. Whiteâs summary of their architectural tendency: âBasically just mostly overengineer.â
He believes existing LLMs already suffice for much empirical biology because even the top 1% of human guessers may perform about like the top quintile or quartile at predicting experiments. Waiting ten years for smarter models might not alter the readiness to automate substantial portions of the method.
His calibration changed more slowly than the technology: âI overestimate the speed of things on month scale and I underestimate things on year scale.â Each month felt disappointing, but 2023â2025 produced enormous progress; FutureHouseâs declared 10-year mission looked much closer after only two years.
4. Nature and laboratory logisticsânot the first hypothesisâare the binding constraints
A hostâs systems-level pushback was that wet-lab work must be the constraint. White agreed: even âwhatever Opus 7 or GPT10â can only propose an initial experiment before it needs information that cannot be computed, because real biological systems contain too much state to simulate exhaustively.
Robin demonstrated the desired loop: an agent proposed an experiment, humans performed it, an agent analyzed the result, and the system proposed what to do next. The core product is therefore not a one-shot oracle but a process that repeatedly earns new information.
Todayâs blocker may be âsomething sillyâ: knowing what reagents are already present, their lead times, the experimentâs cost, and what laboratory capacity is available. Choosing between âGPT 5.2 Codex Max or Opus 4.5â matters less if neither sees the operational context.
This changes the infrastructure requirement. Perfect robotics is not always necessary; an agent can communicate with a CRO or guide a scientist. What matters is reliable state, feedback, and provenance across the physicalâdigital boundary.
5. Scientific taste resisted direct preference learning
White defines scientific taste as judging what is exciting rather than merely correct or feasible. Research topics reflect not only utility but accumulated careers, communities, and preferencesâwhy one organism or mechanism attracts attention while another equally tractable one does not.
After repeated Monday 8 a.m. debates, White and Rodriques tried âthe dumbest thingâ: generate hypotheses, show pairs to people, and ask which they preferredâeffectively human-preference training over scientific ideas.
Raters attended strongly to tone, factual specificity, and whether an experiment looked actionable. They were much worse at assessing the load-bearing question: if the hypothesis proves true or false, how much does the result alter the world model?
Cosmos moves preference downstream toward observable consequences: which report someone downloads, which discovery they choose, or whether an experiment succeeds. When a host challenged its roughly 52% or 55% result, White clarified that the weak score concerned interpretationâwhether a finding was exciting or novelânot raw experimental correctness.
6. Robin shifted Whiteâs trust from expert rankings to verifier-in-the-loop science
Googleâs AI co-scientist impressed White by generating many hypotheses and using tournament-style LLM dialogue to rank them. Robin took a different route: iterate through literature, data analysis, lab context, physical experiments, and another cycle of updated hypotheses.
In dry age-related macular degeneration, specialists broadly agreed on a top 10 but produced noise beyond that. Four weeks of experiments identified a mechanism and a repurposed ROCK-inhibitor drug, likely ripasudil, that had not ranked as the human favorite.
White preserves the novelty hedge: a masterâs thesis may have mentioned the mechanism on page 38, though he suspects it meant wet rather than dry AMD. He conceded one prior report might exist rather than overselling an uncontested discovery.
The experiment changed his mind: literature searches, data analysis, unit tests, and wet-lab results deliver more signal than âwe like this one better.â A host borrowed Max Tegmarkâs phrase ânatureâs computerââthe physical world becomes an indispensable compute cycle inside the agent loop.
7. Enumeration works when filtration and provenance remain cheaper than experiments
Robinâs original ROCK-inhibitor direction largely arose through enumeration. Whiteâs advantage thesis is blunt: âIf you canât be smarter, you can try more times,â then reject candidates using literature, bioinformatics or GWAS evidence, and existing datasets before paying for experiments.
Earlier âtiling treesâ tried to branch through every method, substrate, and choice, but produced nonsensical hypotheses that would waste laboratory capacity. The host argued that LLMs can often filter obvious garbage about as well as experts, though he warned that domain-specific gotchas remain.
Provenance is structural, not cosmetic. PaperQA attaches every sentence to a page; Robin can connect a conclusion to a literature finding and the exact Python line producing an analysis result. That audit trail makes enumeration inspectable rather than an opaque fountain of ideas.
On BixBench, biological-data agents achieve roughly 60â70% correctness, while human analysts agree only about 70% of the time. Running an analysis 100 times can expose consensus and sensitivity to imputation or other choices, separating data noiseâaleatoric uncertaintyâfrom disagreement created by analytical choicesâepistemic uncertainty.
8. Cosmos uses an evolving world model as the shared state of discovery
White describes himself as a âLego guyâ: ChemCrow handled medicinal chemistry, unreleased ProteinCrow protein design, Ether0 chemical intuition, PaperQA literature, and separate agents data analysis and reporting. Robin first assembled these pieces into a concise Python workflow.
Cosmos emerged from asking what Robin was actually updating. Its world model is not merely memory or accumulated papers: it changes over time, accepts inputs, produces predictions, and can be evaluated for calibration.
Early attempts grounded the world model in literature and stalled because literature supplied no genuine experimentâresult cycle. First author Ludo persisted another week or two after the team paused; connecting the data-analysis agent finally let the model explore ideas, observe results, and update itself.
Whiteâs analogy is a git repository: todayâs filesystem distills a long graph of commits, reviews, and contributions into shared working state. Public Cosmos is close to the internal system, while larger versions can run longer, use GPUs, test pre-release models, and use protein or chemistry tools; Cosmos has BoltzGen internally, while external tools can be exposed through APIs.
9. Experimental-data learning beat first-principles simulation in Whiteâs decisive comparison
Whiteâs provocative call is that MD and DFT are overrated, having consumed âan enormous number of PhDs and scientific careers at the altarâ of beautiful simulations. The host separately estimated that, pre-ChatGPT, perhaps 20% of the worldâs computing power went to simulating water; White reacted strongly to the example.
His own quantum water calculation used about 1 million CPU-hours over five months to model proton hopping through water. The payoff was not a de novo discovery but hyperparameters reproducing known effects; he said DFT simulations may use 330 Kelvin when intended to represent room-temperature water to compensate for model error.
The structural problem is that important catalysts contain grain boundaries and dopants and are otherwise complicated, while tractable simulations favor pristine systems. âSimulations simulate really boring things really well. They donât simulate interesting things very well.â
D. E. Shaw Research built bespoke silicon and clusters to test MD-based protein folding at enormous scale. White imagined governments buying perhaps five machines to fold one or two proteins daily; AlphaFold instead learned from X-ray crystallography and ran on a desktop. âThe machine learning on experimental data beat out first-principles simulation by a very large margin.â
10. Natural language remains Whiteâs chosen connective layer, with explicit limits
White still defends âthe future of chemistry is language.â Solubility models, population data, papers, and code need a common interface; humans continually invent words until disparate observations and abstractions can be discussed together.
A host pressed that chemistry also communicates through molecular graphs, geometry, diagrams, and SMILES. White traced the representational ladder from bonds to conformational ensembles, electron density, electron correlation, relativity, and environmental effects: exhaustive fidelity eventually consumes all available compute, so every system must draw a line.
Quantum mechanics supplied the harder objection: perhaps mathematics expresses consequences that words cannot. White conceded that useful scientific language may need equations, SMILES, diagrams, video, or gestureââI donât make sure in our house everything is described with natural languageââwithout abandoning language as the main junction.
His broader method is to adopt strong opinions even when ânot fully correct.â Betting that scientific agents were the future let FutureHouse skip foundation-model detours; optionality can become paralysis. He expects eventually to drop the language thesis, ânot yet though.â
11. Ether0 showed that scientific verifiers become adversarial engineering projects
Ether0 asked whether chemistry could gain verifiable rewards like mathematics or code. A seemingly simple taskâproduce a molecule containing specified counts of nitrogen, oxygen, and hydrogenâbecame a catalogue of ways a model could satisfy the checker while ignoring chemical usefulness.
The model repeatedly chained implausible numbers of nitrogens. White insisted six-nitrogen compounds were impossible, only for a Nature cover to report humanityâs difficult synthesis of one in 2024 or 2025; Ether0âs outputs were still unsynthesizable reward hacks, not prescient discoveries.
Requiring purchasable reagents triggered another exploit chain: remove one atom from the target and âbuyâ the rest; add purchasable nitrogen that does nothing; then add an acid that participates only by moving one atom. White ended up building a purchasable-compound catalogue and Bloom filter, asking, âWhy am I doing this? How did I get here?â
The team used GRPO with modifications including DAPO and special clipping, yet mundane leakage still dominated. One model failed because training reagents were alphabetically sorted while test reagents were notâthe learned strategy exploited ordering rather than chemistry. âBulletproofâ verifiers proved far harder than supervised pre-training.
12. Automationâs safety and labor boundaries remain open questions
Whiteâs 2023 assessment of chemical, biological, radiological, and nuclear risk was that dangerous targets and synthesis routes were often already public; expertise and material handling, not missing facts, remained the constraint. Tacit protocols and scale-up troubleshooting were more credible concerns, prompting labs to test and filter such requests.
White acknowledged possible second-order logistical effectsâfinding centrifuge vendors, estimating prices, or navigating KYCâbut said AI had not meaningfully accelerated the core work in practice. The host proposed a possible second wave involving real-time assistance or computational scenarios that seemed too remote two years earlier; White called that framing vague. The host also said his own safety opinion was not fully formed.
On employment, White invokes Jevons paradox: science has no finite stock of 100 remaining discoveries, so cheaper discovery could expand demand. Scientists become âagent wranglersâ or âCosmos wranglers,â exploring 10Ă or 100Ă more ideas simultaneously.
He nevertheless concedes friction: a pharma or materials CEO may spend another $1 million on an AI scientist instead of hiring ten people. When a host asked why humans must remain, White returned to taste and science as something humans appreciateâthen admitted, âMaybe youâre right. Maybe there is no point for humans.â
Full transcript
MD was supposed to be the protein-folding solution. There is a great counterexample: a group called D. E. Shaw Research. They had similar funding to DeepMindâprobably more, actually. They tested the hypothesis to death that molecular dynamics could fold proteins. They built their own silicon and their own clusters, and had them all taped out themselves. They burned the algorithms to run molecular dynamics into the silicon. They ran molecular dynamics at huge speeds and huge scales.
I remember David Shaw came to a conference once on molecular dynamics. He flew in by helicopter and was this pretty famous, kind of rich guy. He gave an amazing presentation about the special computers and the special room outside of Times Square, and what they could do with it. It was beautiful and amazing. I always thought protein folding would be solved by them, but it would require a special machine. Maybe the government would buy 5 of these things, and we could fold maybe 1 protein a day or 2 proteins a day.
When AlphaFold came out and it was like, you can do it in Google Colab, on a GPU, or on a desktop, it was so mind-blowing. I forget that protein folding was solved. I always thought that was inevitable, but the fact that it was solved and you could do it on your desktop completely floored me. It changed everything.
This is the first episode of the new AI for Science podcast on the Latent Space network. I'm Brandon. I work on RNA therapeutics using machine learning at Atomic AI.
My name is R. J. Haniki. I'm the co-founder of Mira Omics, where we build spatial transcriptomics AI models.
The point of this podcast is to bring together AI engineers and scientists, or bring together the 2 communities. These are 2 communities that have developed independently for quite some time, but there have been attempts to combine them. Only now, after many years, are we starting to see some of the big developments play out in the real world and start to solve key scientific problems.
There's no one-size-fits-all solution. You need domain expertise. You need people on both sides of the aisle who can really talk to each other, work together, and understand both the modeling and all of the real subtleties of the system you're actually trying to work on. We hope that we can connect these communities and provide a starting point for this new era of AI and science to move forward.
So, without further ado, let's get started on the first podcast. We're really happy to have in the studio today Andrew White, co-founder of FutureHouse and the newly formed startup Edison Scientific. Rather than introduce him, I'll let him introduce himself.
Hi, I'm Andrew from San Francisco, a former professor now running 2 startups: 1 that's a nonprofit research lab and 1 that's a for-profit, venture-backed company. We're trying to automate science.
We're going to get into all those points.
Yeah, I'm really happy to be here. Thanks for having me on.
I personally want to know about the jump from academia to industry, or quasi-industry. I would love to hear that story.
Yes. I guess that's the whole story, right? I did my PhD at the University of Washington, and I worked in a group with, I think, 19 people doing experiments and 2 people doing simulations. I was working on a topic called molecular dynamics, which I think is suddenly becoming interesting again as everyone's looking for ways to generate data from first-principles simulations.
Molecular dynamics covers basically everything involving molecules moving around in dynamic systemsâbiology, things like that. The complement in materials science is density functional theory, where you can model chemical reactions and solid systems. I was working on that, and we worked on biomaterials.
The goal of my PhD was trying to find what are called non-fouling materials. In biological systems, whenever you put a foreign object into the body, it will trigger a response. That response, called the foreign body response, basically encapsulates it in a layer of collagen.
This is actually exploited for some implants. If you get a pacemaker installed, the body coats it with this collagen, so that if you go to change the battery, you can almost change the battery out without even bleeding, because the body has completely encased it. This is great for pacemakers, but for a glucose sensor or a brain-computer interface, or BCI, that's not so great. That's why some of those things have a limited lifetime: eventually, your body treats them as a wound and heals andâ
Rejects.
Yeah. Some rejection is immune-based. If the body can see anything on itâif it can see some ligand that it can bind to with antibodiesâthen you get this inflammation, which is a rejection response you see in organ transplants. With materials, the body just goes, âOh, there's a wound or something here,â and covers it up.
Okay.
I think the research in that field has gone on for a long time since I left my PhD, and there were a lot of theories about how it was related to the mechanical properties of the materialâwhether it was spongy, or whether it was trabecular, with a bunch of little pores in it. We worked on the theory that it had to do with how hydrophilic the material was.
I was the only one working on computers in this group. I couldn't figure out how to connect what's on the computer with what's done in the lab, because you can make a simulation of, say, 10,000 particles or 10,000 atoms, and it's like, well, this is not going to model a human body and an implant. That's a lot more atoms involved.
I had a good time. We did some cool biomaterials work, and I learned a lot. But then, when I did my postdoc, I thought, âOkay, we're going to try to merge experiments and simulations.â I worked on this theory called maximum entropy. It's about how you take complex simulations and match them to limited observations.
It's like the inverse of machine learning. Machine learning is like you have simple models and you're mapping them to a lot of data, whereas I had complicated models and was trying to fit them to very little data. It was fine. It was great. We wrote some papers, and it was useful.
Then I started my research group at the University of Rochester, applying these methods to model peptides.
Yeah.
I'm always too early for things. We studied peptides for 4 or 5 years, and it was a cool niche fieldânot that popular. Now peptides are the hottest thing ever. I think there's even a peptide rave, I heard about a couple of weeks ago.
When I was an assistant professor, nobody cared about peptides. We worked a lot on different ways to combine them. We looked at different experimental methods that we could use with these molecular dynamics simulations of peptides.
Then, in 2019, I was out on sabbatical at UCLA. They have a place called the Institute for Pure and Applied Mathematics, which is an institute where people can go and do a sabbatical and learn new methods. They happened to be doing machine learning for physics. I think the name of the program was something like âMachine Learning for Physics and the Physics of Machine Learning.â
Okay.
It was a cool concept. Yann LeCun was there, and Frank Noé was there, who's a big guy in Europe in this field. Terence Tao came by. It was a great group, and everyone was jamming. It was 2019, so there hadn't really been the big hype, especially in non-computer-science fields.
Right.
I came back from that and thought, âWell, I've got to teach a class on this.â I wrote a book about how you can apply these methods in chemistry. It was a very niche field because every machine-learning class that my PhD students could take at the timeâthis was when I was a professor at the University of Rochesterâwould always end with, âOkay, this is an RNN, and this is what you need to know,â or, âThis is how you do image classification.â
In chemistry, it's all about graphs, right? It's all about how you represent these graph structures. It's all about symmetry and geometry. That wasn't very popular at the time, but you had Max Welling on before, and he's the godfather of geometric deep learning.
Yeah.
So I wrote this textbook about these methods. There was a bunch of interesting mathematics to it. I had a good time with it.
Then I think I was following the news in the space, and the original Codex came out. I had been looking at transformers for a while and just tinkered with them. We started trying them on chemistry tasks, and we were really impressed. We wrote a benchmark. This was around 2019 or something. We wrote a benchmark of verifiable rewards in 2019âmaybe it was 2020 by thenâbut we were a little ahead of the curve.
Ahead of the curve.
A little ahead of the curve again, yeah. Here's a function, and there's a task: I have the body of a function for a Markov chain Monte Carlo simulation, it's missing some pieces, and you complete it. Then we had a verifier that would see whether it was a valid MCMC simulation.
We wrote this paper, and it ended up coming out, I think, in 2021 or 2022, because it took a long time to bank enough questions. I wrote an opinion piece about how transformers could change how we think about chemistry and how we teach it.
Then OpenAIâsome people there. Lama was there.
She saw this paper and reached out, saying, âHey, weâre building this new model, and we think itâd be great to red-team it to see what could happen with these models if theyâre applied to chemistry or biology.â I was a red teamer for GPT-4, and I was using it for 9 months or something before releaseâit was in August. So GPT-4 came out in March, and I was using it in August.
Yeah.
Then the ReAct and MRKL papers came out. I think Shunyu Yao wrote the ReAct paper, and I plugged it into GPT-4 in the fall. I was like, âWow, thereâs so much stuff coming out.â
With ReAct.
Yeah, and it was really exciting. Then, when GPT-4 came out, I released this paper called ChemCrow. I worked with Philippe Schwaller in Switzerland on this, along with IBM.
So that was ReAct applied to chemistry.
Yeah. What we had was a cloud lab that IBM built in Switzerland. We had GPT-4 operating the cloud lab, and I had written a literature-research agent that did agentic RAG. Again, nobody knew what agentic RAG was at the time. I think Harrison Chase had written a blog post about some ideas there, so I stole some of his ideas. Heâs a really smart guy.
Basically, we applied that, and we saw some really cool stuff. It was really exciting, and then we wrote the paper. It set off this crazy storm where everyone was having a lot of anxiety about AI progress.
Yeah.
I ended up visiting the White House. I guess my paper was the only time a preprint or peer-reviewed paper was presented to the president on their schedule, in a 30-minute block.
Wow.
The national security adviser at the time, Jake SullivanâGod, I was confused aboutâ
Yeah, no, sorry. One of them is a talk-show host, and one of them is the national security adviser. I forget which is which.
That guy.
That guy.
Yeah. He had a presentation about our paper, and they presented it to everyone because there was a big tech CEO summit at the time. They sent Sam Altman and some other CEOs out there, and theyâ
Was this âThe Future of Chemistry as a Language,â or a different one?
This was the ChemCrow paper. Sorry, I probably should name these things.
It was crazy. They had me go out there, and I met a lot of three-letter agencies I didnât really want to meet. Somebody from one of the three-letter agencies asked, âHow does this change explosives?â The three-letter agencies were asking, âHow does it change breakout time for nuclear-weapons research?â
I was like, âI donât know. Iâm not really sure.â But it turns out that there arenât that many people who are world experts on AI and science, right?
So whatâs the answer?
Yeah, I agree. Good question. Weâll come back to that.
Okay, yeah, letâs come back to that.
In the end, I had a lot of energy and excitement about this area, so I took a sabbatical from the University of Rochester. It was Sam Rodriques, and Sam had been talking to Eric Schmidt and Tom Kalil, who was also at the National Security Council in the Obama administration, about how to scale up these ideas.
Sam had this concept of focused research organizations: How do you do science not in academia, and not in one of these near-monopoly tech companies, these big labs? I thought, âHey, we should do this around agents for science, or AI for science.â
I love Sam. He pushes me to come up with really lofty ambitions. We decided to make automating science the goal, instead of seeing what fun stuff we could do with agents in science. I think that was maybe the real mission. Of course, automating science is the long-term mission.
Yes.
That was what led to FutureHouse.
That was very long-winded.
Yeah, no, no, thatâs great. So you chose to leave a tenure-track position?
I was on sabbatical, which is a beautiful concept. But then I did resign my tenured position when we co-founded Edison. I had been on sabbatical for a very long period of time, and at a certain point I just had to resign my tenure. I resigned my tenured position in June.
Oh, so thatâs only recently.
Yeah, only recently.
And you just felt like this is the direction of your career?
Yeah. I got tenure, and I had these early-career awards, like the NSF CAREER Award. It was great, and I think academia is really exciting. But I thought that, right now, this kind of area is difficult to do in academia, and itâs so exciting that I think you can take bigger bets.
Having a tenure position and writing research grants is maybe not the biggest bet you can take on a field.
Yeah. So now we have a venture-backed startup called Edison, right?
Which we spun out of FutureHouse. We took a lot of the ideas and weâre trying to do this at an even bigger scale right now.
Yeah.
Edison was always kind of the plan, going back to Samâs idea of a FRO, or focused research organization. He always had this goal of doing fundamental research in a tightly scoped nonprofit that could explore, and then you would have that as a natural arm for spinning things off.
Yeah. You know, venture-backed.
Yeah, I think thatâs right. I think some things that make that not as clean these days are how expensive AI research is and how expensive GPUs are. I donât think we can repeat it many times from FutureHouse. It might be an end-of-one thing right now. It just may not be. I donât knowâif venture capital keeps growing, then maybe we can.
But I think we took a lot of the ideas from FutureHouse. Another thing is that I think we expected it to be harder to automate science. Actually, itâs really hard. I feel like Iâm always miscalibrated in this domain, but itâs always hard to predict progress.
Yeah.
I overestimate the speed of things on a month scale, and I underestimate things on a year scale. The 2 years from 2023 to 2025 represented an enormous amount of progress. It always felt like things werenât going as fast as I thought, but when you look back on it, you realize, âWow, thereâs been a lot of progress.â
I think that in FutureHouseâand Sam actually regrets us writing thisâin the original marketing, or the announcement, it said that it was our 10-year mission to automate science. Now itâs like, âOkay, yeah, 2 years later we had Cosmos, and things are going so much faster.â
Thereâs also something you notice in San Francisco: Itâs actually kind of hard to find problems that are hard enough to be a challenge for language models but not so hard that theyâre impossible. Thereâs this gray zone, and I feel like thatâs where we are right now.
We can automate so much of the scientific method because it turns out, especially in a field like biology, which is very empirical and limited, that the top 1% guesser of what will happen in an experiment and the top quintile or quartile are about equal. Even if we wait 10 years and get even smarter models, I donât think itâs going to change the fact that weâre ready to automate a lot of science with existing LLMs.
What do you mean by âautomate scienceâ? Thatâs a pretty loaded statement. There are lots of ways of thinking about that.
We try to draw a line between groups that are trying to model something, like the cell, how proteins fold, how antibodies can be designed, or maybe virtual cells as an example. If theyâre trying to use machine learning or AI to model some very specific system, weâre trying to automate the cognitive process of scientific discovery: making hypotheses, choosing experiments to do, analyzing the results of experiments, and using those results to update your hypothesis or your confidence in those hypotheses.
That leads to a world model of, âOkay, this is how I understand this process to be,â and that begets new hypotheses or new experiments. We want to automate that sort of loop.
We thought that we would have to build up a whole new organization from the ground up for agents. That means automated labs, putting all the papers in one spot, and getting APIs wrapped around everything. But over time, the models have gotten better and better, so we had to stop and rethink: We donât actually have to hold their hands so much anymore.
They donât necessarily need to have an automated lab. They can write an email to a CRO, or they can tell you what experiment to do, and you can take a video of yourself doing it and show it to the model. The model can say, âOkay, well, this is what happened.â
Itâs been a really interesting experience. Sometimes we overengineer things, and sometimesâactually, basically, we mostly overengineer.
I always think about systems, and science is a system. I think about the scientific process as a system in terms of constraints: What is the bottleneck in the system? So what is your hypothesis about this?
Not knowing a ton, in my mind the constraint of the scientific process is the work you do in the lab, and thatâs notably missing fromâwell, not entirely missing fromâyou mentioned automating the lab and everything. How are you thinking about this?
Yeah, I think youâre right. Basically, the best modelâwhatever, Opus 7 or GPT-10âreally can only propose the first experiment, maybe a slightly more clever one. At a certain point, you just need information.
There are some little calculations you can do, but there are more atoms in the brain than you could ever simulate, even if you had all the energy from the sun. I think you could simulate maybe 1,000 brains in real time with all the energy in the sun, because thereâs just too much information.
Science really hits these bottlenecks where you actually have to go measure things.
Yeah. We definitely think about lab-in-the-loop situations. One of our papers, which was called Robin, had one of our agents propose an experiment. We did the experiment, and then we had our agent analyze the experiment, propose the next experiment, and continue that kind of loop. I think that's where you want to get to.
Yeah. So what is the bottleneck in that?
I don't think it's the intelligence of the first experiment. I think the bottleneck might be something silly, like knowing the lead time on all the reagents that you need and what is available in the lab, right?
Yeah, yeah, yeah.
I think whether GPT-5.2 Codex Max or Opus 4.5 is going to do better probably doesn't matter. It's just a matter of which one is going to have all the information about what's in the lab, how much it will cost, and how long it will take.
Right?
And also, I guess, the kind of frontier that I think about for these models is taste, which is a broad category. A lot of science, of course, is about accelerating technology, improving the economy, improving people's life expectancies, and making everyone happier. But a lot of what is done in science is based around human preferences.
Why do people study a particular worm? There is a theory that studying the worm has led to good medicines or to discovering new genes. But people also studied it in the past, people's careers depend on that worm, and people want to write papers about that worm. There is a human element to some of this, and I don't think these models capture that very well: knowing what is an exciting result and what is a boring result.
I see.
So I think that's scientific taste. It's a broad category of all these things.
How do you define taste? I know I have some fun anecdotes about this, but I'd like to hear what you thought.
Yeah. We actually sat on this idea and argued about it for a long time. Sam and I usually meet every Monday morning at 8:00, and we're both caffeinated and ready to argue about stuff like this. We had a lot of Mondays where we talked about scientific taste.
In the end, we said, âOkay, let's just do the dumbest thing,â which was to have our agents make hypotheses, put them in front of humans, and have people say, âI like this oneâ or âI like that one.â So we just did RLHF on hypotheses, and we learned a lot about how bad RLHF is.
People really paid attention to the tone, the details, and how many specific facts or figures were in the hypothesis. They paid attention to actionabilityâwhether the experiment was feasibleâbut what people didn't really pay attention to was, I don't know how to describe this, if the hypothesis is true, how does it change the world? If the hypothesis is false, how does it change the world? It's how much information you gain. It's not really information, but impact or something. That really didn't come through from those tests.
We said, âOkay, well, this is maybe one strategy,â and went back to think about it more. We then took a pause from that research and made Cosmos. Cosmos has taste baked into it. At the end of the day, there will be some report, and we're working on generalizing this. At the end, I can say, âI made these discoveries,â and a person can say, âGreat, I want to download that one,â or, âI like that one,â or, âI don't like this one.â That rolls up to some hypothesis that came earlier in the process, so we think we can get to end-to-end human preferences.
So you mean the feedback loop is the click?
It could be the click. It could also be that we do an experiment. Sometimes in Cosmos, you can ask to end an experiment, and then go see whether the experiment was a success or failure, or something like that.
I guess we've brought it out of this hard-to-quantify question of whether this is a good hypothesis or a bad hypothesis and into something where you can see the downstream consequences of the hypothesis. Humans have a very strongly calibrated nose for science. Maybe you could argue that there are sociological effects across the community, but ultimately, good scientists often know right off the bat whether something is likely to be useful or not.
How many attempts did it take before you started to see results that seemed useful to you? You've been working on this for, I guess, 2 years now.
I think when the AI co-scientist paper came out from Google, it was a really interesting idea to do this tournament-style, or just pairwise ranking, of hypotheses. I think AI co-scientist is a very interesting counterexample to what we built.
What we built is something with either lab-in-the-loop, data analysis-in-the-loop, or literature research-in-the-loop, where you're iterating on an idea. I think AI co-scientist took a very different approach: âLet's list all the ideas and then try to come up with a filtration process to find the best hypothesis.â
AI co-scientist produces these very long reports where it says, âWe really tested this idea,â with lots of dialogue, and it was very interesting stuff. I was really impressed with the paper that came out. Then we had this Robin paper, and one of the things that came out of the Robin paper was that the hypothesis people thought was best was not the one that led to success in that paper.
Interesting.
It was in age-related macular degeneration, or AMD. Basically, part of the eye is going blind because you have this accumulation of drusen in the eye and can't clear it out.
That's the major cause of blindness in people over 60. Ollie, who works on the Hillâ
Yeah, yeah. He'll cringe when he hears me say that, butâ
Something like that.
Something like that. Sorry, Ollie. In that one, we went to optometristsâor ophthalmologists; I get those confused as well. Sorry, Ollieâand essentially asked them which hypotheses they thought were good hypotheses, which they thought would lead to a good mechanism for treating dry AMD.
Yeah.
They agreed on the top 10, but beyond that it was kind of noise. Then what we found was that ripasudil was a very good medicine, and it had a mechanism that I think is novel, although there was lots of debate on X. I think there was a master's thesis that proposed this mechanism on page 38. I actually think it was a typo; I think they meant wet AMD. But anyway, I won't belabor the point. I will concede that maybe there was one reported example of it in the past.
That was a really eye-opening experience for me. It was the first really serious test where we went to the lab and spent about 4 weeks on a battery of experiments to see which hypothesis led to a good mechanism and a good repurposed drug.
Right.
It was not as correlated with human opinions as I expected.
Yeah, yeah.
Since then, I think I have a lot more faith in verifier-in-the-loop scenarios, where you have either data analysis, literature search, or you're running a unit test, or you're going and running the experiment. Anything like that is going to give you a higher signal than the vagaries of, âThis is a higher opinion,â or, âWe like this one better.â
Yeah. Max Tegmark called it nature's computer.
Yeah. It's like you have this computer cycle you're running, and nature is part of that computational cycle.
I'm curious. You said that there is a paper that maybe could have proposed where this molecule came from, but do you have some way of interpreting or understanding where that hypothesis originated in the absence of that? Is there a little thought train?
Yeah, yeah, yeah. This is something we pay really close attention to at FutureHouse and at Edison: provenance of information.
Our first sort of agent was PaperQA. Sorry about the name. PaperQA sounds like an email address, but that was an agent.
It really does.
Yeah. PaperQA has every sentence that it outputs accompanied by a citation to a page, so there's a lot of provenance. We basically built everything around that philosophy.
Robin, which is the name of this workflowâor something; you can call it thatâthat led to the result of ripasudil being a good therapeutic for dry AMD has data analysis that shows you which line of Python code led to the result here. Then that goes to another model, which says, âBased on this literature finding and this result from the data analysis, I believe this is the right thing.â
But where does the original idea come from? Going after these ROCK inhibitorsâthe mechanism for the target was basically enumeration. If you can't be smarter, you can try more times, of course. I think that was the theory of the Robin paper: we can put out a whole bunch of hypotheses and then filter them, just like I think AI co-scientist did. You go through a filtration process, but the difference is that in AI co-scientist, the filtration process was other LLMs ranking it with rubrics or personas, whereas our filtration process was literature search and data analysis.
Here's some data.
Is it consistent with the data? Go see if anyoneâs discovered it in the literature or if theyâve disproven it. And I think thatâs the easy way to succeed in AI over humans: You can try more ideas faster.
Something Iâve heard people say, and maybe Iâve experienced this in my own life, is that sometimes hypotheses are kind of cheap, especially in biology. In many ways, itâs actually easy to come up with what you think could be happening. And it seems to me that verifying is often a big bottleneckâmaybe the biggest bottleneck. If you have lots of hypotheses and it costs 1/100th of your runway to test each one of them or something, you donât have any shots on goal.
Yeah.
Yeah. So how do you make sure that you are actually enriching for good hypotheses?
Literature and data analysis, right? There was a time when we used something called tiling trees. A tiling tree is a literal brute-force method invented by Ed Boydenâs PhD advisor, and basically the idea is: âOkay, I want to accomplish X. I could try these methods.â Once you pick, âIâm going to try this method,â then you split into 2 different paths: âIâm going to use this methodâ or ânot use this method.â If youâre using this method, you need to have some kind of substrate. âIâm going to try this substrate, or this substrate, or this substrate,â right?
You can basically try to tile the space of all possibilities. We tried some early experiments there, and youâre right: You run into this thing where some of the hypotheses come out and just donât make any sense, and youâre going to waste a ton of effort if you actually test them all. Nowadays, I actually would argue that if you go to an LLM and ask it to evaluate hypotheses, including some garbage ones, it will probably do as good a job as an expert in the field at filtering them out. Thatâs not always the case.
swyx
Yeah, Iâve actually seen that myself.
Alessio Fanelli
Yeah. But there are a lot of gotchas, and I think people can miss those, but I think theyâre actually pretty good. And so Iâm not as worried about hypotheses that can fail fast by an expert looking at them.
I think now the filtration process really happens in literature. And I think the filtration process happens in looking at bioinformatics data, or what we know from GWAS, or other sources of existing dataâas much as you can draw upon.
swyx
Yeah. So with regards to existing data, another maybe contrarian take is that oftentimes the hardest part is just understanding the context of data, where it comes from, and how you interpret it. I can also think from my own life of multiple cases where the data, in some sense, was there, and you had 2 people who were both experts and very smart people who looked at it and drew very different interpretations. In fact, when we were interviewing Heather Kulik, she had some fun stories about using LLMs, and she would find that there would be raw data in a paper that wouldnât agree with the conclusions of the actual paper. And itâs straight from the paper; itâs not even cross-paper talk or something.
Man, Iâm going to be a really boring interviewer and be like, âYes, youâre right.â You know, this is a hard question.
Alessio Fanelli
I think, to give you something concrete, we have a bioinformatics benchmark we call BixBench. BixBench is something we put out, and weâve updated it a few times. Itâs in some frontier LLMsâ system cards; when they release their system card, theyâll mention BixBench. Itâs one of the things they test on.
swyx
Yeah.
Alessio Fanelli
And weâre getting to 60%â70% correctness on BixBench, and we found that weâre actually at the point where humans disagree at this level. Humans only agree on 70% of the analysis. And so itâs true that, when it comes to analyzing data, humans do not agree 100% of the time. Thereâs a certain amount of choice that goes into it.
We try toâso Edison is a for-profit company. Maybe weâre trying to sell some of this stuff to companies, and weâll go to some companies and theyâll say, âOh, we never impute data. Imputing data is bad,â or whatever. And weâll say, âOkay, well, weâll have to change our agent so we donât impute data with them.â But then some other companies are like, âOh, yeah, we impute data. It makes everything easier,â right?
And you want to know what the real modern dark arts areâthat AI-resistant area of the world? Itâs medicinal chemistry. That is the spot where thereâs so much superstitionâ
swyx
Oh, yeah. Everyone is pseudo-religious.
Alessio Fanelli
Yeah, exactly. But you have to be to survive. Otherwise, you get burned out.
swyx
But the religions never agree, either. 2 medicinal chemists will have completely different viewpoints about a functional group.
Alessio Fanelli
Yes, exactly. And I remember talking to somebody who worked at a CRO, and they were like, âOh, whenever company X orders anything, we never put boron on any of the compounds because they hate boron. There was one program that was killed because there was a boron somewhere in the core, and it led to some toxic side effect. So no boron for this company.â This company, they love things to be fluorinated or something because they think itâs great for the ADME properties, right?
And so thereâs all this stuff where you reach the point whereâI donât knowâhuman-bias level or human-disagreement level, and I think weâre getting to that point in data analysis. And so, of course, you will see that if I take the raw data from a paper and analyze it myself, I will get a different conclusion.
One of the cool tricks you can do, going back to this brute-force thing, is that I can go to our agent and run it 100 times and take the consensus analysis. Or I can say, âEven if you make these 3 different choices in your data analysis, you get the same conclusion,â right? Or, âThis conclusion is somehow sensitive to those choices.â Then you can say thereâs even terms like epistemic versus aleatoric uncertainty. Itâs like, âThis is aleatoric,â which means, âI think itâs noise from the data,â or, âThis is epistemic uncertainty,â which means, âI think there are some choices being made. There are some differences that lead to the disagreement.â
Anyway, thereâs a Donald Rumsfeld formulation of this as well: the known unknowns. And, yeah, the aleatoric-epistemic debate there.
swyx
Interesting. This kind of digs into your Cosmos a little bit. I glanced at the paper, and one of the things that jumps out is that there was a certain class of problems for which it was only 50-some percent accurate. Can you talk a little bit about that? If Iâm just getting 50% accurate answers and then going into the wet lab saying, âOkay, try this,â only to realize, âAh, the stupid thing told me to do something dumb,â how do you handle that?
Alessio Fanelli
I would say, first of all, that 50% is actually pretty good, because itâs rare that experiments in the lab are actually coin tosses, right? There are usually a lot more outcomes than binary.
swyx
Yeah. Yeah. Sure. Okay.
Alessio Fanelli
But that particular number was human agreement in the interpretation of the results. We asked people to evaluate different aspects of Cosmos. We had them evaluate the data-analysis decisions, and we asked people to evaluate the literature: âDo you agree with its finding in the literature?â That numberâthat 50%âcame from Cosmosâs interpretation of some of the analysis.
So it might go into the literature and find this result, and then say, âWow, this is super exciting. This is amazing.â Or it might do data analysis and say, âThis is a novel discovery. Really excited about it.â And then people would disagree: âThatâs actually not interesting,â or, âI donât agree with the interpretation of it.â
swyx
So itâs like picking bad problems, maybe.
Alessio Fanelli
Yeah, in the negative class. And so I think that 52% or 55%, whatever it is, thatâs interpretation. And so I agree: I think thatâs where, like I was saying, the frontier right now is scientific taste.
And so thatâs what weâre working on right now: How do you get that interpretation to match?
swyx
You step back and just introduce Cosmos from a high level. Iâd actually be even curious to hear, starting from ChemCrowâand, you know, you have PaperQA, Aviary, Ether0âIâd like to hear a little bit of the lineage and how those different decisions were made. What were the key learnings, and how did you get to where you are now?
Alessio Fanelli
Yeah. I could retcon and tell a really great story about how we arrived at Cosmos, but I will say that, to a large extent, we just try a lot of stuff. Sometimes it works, and sometimes it doesnât.
Iâll say that weâre veryâIâm a builder. I like to build things piece by piece. Iâm probably some fancy word for it, but Iâm a Lego guy or something. My vision was that we would make an agent that does this part of the scientific process, an agent that does that part of the scientific process, whatever.
And so we had ChemCrow, which was going to help us with setting up our medicinal chemistry work. We had ProteinCrow, which we havenât released. I donât know if we will ever release it, but ProteinCrow is for designing proteins we might need for some part of our workflows.
swyx
Or we had a data analysis agent. Itâs an agent: an LLM plus tools.
Alessio Fanelli
Okay.
swyx
Ether0 was, like, âOkay, we noticed that frontier models canât work with molecules very well, so letâs make a model with intuition for medicinal chemistry.â That was what led to Ether0. But then Sam really pushed us: âLetâs just do the whole thing. Letâs just try to build an AI scientist. Letâs just try the whole thing.â
That was what led to Robin. Robin was, âLetâs just take these agents we already have and put them in a workflow.â Basically, you could express it in a concise Python file: try a whole bunch of ideas, then go see if they all filter through the literature or if theyâve been disproven, and then come up with experiments that you could do in a wet lab.
Alessio Fanelli
Yeah.
swyx
This is our inventory list. Then go analyze all the data, go back, and repeat the process. Thatâs what Robin was.
Then we came across Cosmos. We were trying to understand what process Robin was automating, and it came from this idea of a world model. When we first started Edison, we were thinking, âWhat do we want to change about this? What is new here?â
We spent some time thinking about the scientific process: What is actually going on in my brain? I have some understanding of the world or the phenomena Iâve studied, and thatâs my world model. A lot of the actions I take are about trying to update that world model. Itâs something that changes over time, but itâs also practical: I can use it to make predictions. I know from this experiment this will happen. Thatâs why itâs a model and not just memory, or a bunch of papers or something like that. Itâs supposed to operate.
In Cosmos, we tried this idea out. Ludo, who was the first author on the paper, tried a whole bunch of ideas around world models, and we kind of thought they werenât really appropriate. We tried a lot of different ways to do thisâMethod A, Method B, Method Câand they were okay. So we all decided to take a break.
Ludoâs project didnât work on trying to do this world-model stuff. He was like, âIâm going to keep trying it.â Ludo is a very stubborn person. So he tried it for, I donât know, a week or 2 weeks, and he was quietly like, âHey, can you guys come take a look at this?â
We were like, âWow, this is actually really cool,â and then we started building on it and jamming, really. I think what Ludo figured out is that you have to get this experiment-loop thing. You have to let it run, and the data-analysis agent is what got us in the loop.
If you put that in the loop, it can really update this world model, because we were trying to build it around literature before. When you build it around literature, there arenât really experiments you can do and then see the results for. That was our surrogate: literature. It just wasnât working. Data analysis actually really lets you explore ideas, and so that was what led to Cosmos.
In Cosmos, we basically had all the pieces sitting around. We were working on world models, a data-analysis agent, and a literature agent. We had built a platform for scientific agents, too, so we had things that could write a LaTeX report and things that could make nice plots. Then we put that all together, and a world model was sort of the glue that allowed it to fit together. Yeah.
**swyx**
An analogy is, in coding agents, GitHub is sort of the glue. Thereâs some shared repo and everyone works on the repo. Software engineers have spent lots of brain cycles thinking about how to coordinate and organize working on code together for a long time.
So the world model is actually like a memory system, kind of.
**Alessio Fanelli**
Yeah, you can think of it as a memory system. We think about it as a model, so you can put in input and it will output predictions, and we think about calibration.
But really, it is a big bundle of information that we accumulate over time, distilled in some way, and that is what allows us to do this. You can think about a GitHub repo as a distillation. Really, thereâs a long graph of commits that lead up to it, and the current file system in that Git repoâ
I keep saying GitHub. Iâm such a corporate shill here. Get your Git repoâ
Itâs a distillation of all the work that people have put into the pull requests and the commits. I think thereâs a nice analogy between a Git repo and what a world model is.
**swyx**
I see.
**Alessio Fanelli**
And I think thatâs just what allows us to automate scientific discovery so well.
**swyx**
Can you talk about how you implement a world model, or is that sort of secret sauce?
**Alessio Fanelli**
Thatâs our secret sauce right now, you know?
**swyx**
Thatâs fine.
**Alessio Fanelli**
Yeah, no, itâs fine. People have asked around.
**swyx**
One thing thatâs notably missing is the simulation, right? Dynamics, or Boltz, orâ
**Alessio Fanelli**
Yeah, I want to help you guys pump up your views here. I think molecular dynamics is overrated. In factâ
**swyx**
Coming from someone who goes in the thumbnail, you know.
**Alessio Fanelli**
Yeah. And DFT is overrated. In fact, DFT may be even more overrated than the numerics. I think these methodsâ
**swyx**
For materials or for biology, or for both?
**Alessio Fanelli**
For materials.
**swyx**
Okay.
**Alessio Fanelli**
And I can explain more about that. Basically, MD and DFT have consumed an enormous number of PhDs and scientific careers at the altar of the beauty of the simulation.
**swyx**
Also, random interjection: I did an estimate once. I think, pre-ChatGPT, something like 20% of the worldâs computing power just went to simulating water.
**Alessio Fanelli**
Oh my God, water.
**swyx**
Yeah.
**Alessio Fanelli**
I had to deal with so many water simulations. I did DFT simulations of water, and they are so annoying. I used these big computers from the Department of Defense, and I spent, I donât know, 5 monthsâand, by the way, in the pre-training days, 5 months of compute is actually a really long timeâsimulating water with quantum effects and a grotesque mechanism for how a proton hops through water.
Itâs on YouTube. Itâs my number-one YouTube video, and it represents, I donât know, 1 million CPU hours of compute. It was one of the biggest computations that Iâve probably done in my life so far. Maybe Ether0 is bigger, but it took a lot more work.
**swyx**
And whatâs the point? What did you learn?
**Alessio Fanelli**
All I learned was which set of hyperparameters reproduces some physical effects of water. But none of it was de novo, right? This is the issue with molecular dynamics and DFT: They donât model the world correctly.
So we have to invent little stories we tell ourselves, like, âWeâre making good inductive biases,â and then it models the world more correctly. In DFT, you simulate water at 330 Kelvin when you want room-temperature water.
**swyx**
Is room temperature 330 Kelvin?
**Alessio Fanelli**
No, itâs not. Thatâs a little too hot, right? The issue is that people just make up these things. Or, I donât know, GGA, BLYP, or B3LYPâall these different methods are clearly empirical, and then they bolt them onto DFT and say, âLook, itâs a first-principles method.â
But actually, you made a whole bunch of choices and overfit to the validation data to get this to work. I think MD and DFT are like that because if you go look at the catalystsâwhat catalysts change the world? None of them are single-crystal materials that are really well suited for DFT. They always have grain boundaries, they have dopants, theyâre complicated, right? You never capture that with DFT.
I think this is one of the fundamental dichotomies of the world: Simulations simulate really boring things really well. They donât simulate interesting things very well. Thatâs why I donât do DFT and MD anymore.
**swyx**
What about machine-learning stuff like AlphaFold?
**Alessio Fanelli**
AlphaFold was trained on X-ray crystallography data. I think this is the story of MD: MD was supposed to be the protein-folding solution.
Thereâs a great counterexample. The counterfactual, basically, is a group called D. E. Shaw Research. They had similar funding to DeepMind, probably more, actually. They tested the hypothesis to death that MD could fold proteins.
They built their own silicon. They built their own clusters. They had them taped out themselves. They burned the algorithms into the silicon to run MD. They ran MD at huge speeds and huge scales.
**swyx**
Yeah. I remember David E. Shaw came to a conference on MD once. He flew in by helicopter and was this pretty famous, kind of rich guy.
**Alessio Fanelli**
And he gave an amazing presentation about the special computers and the special room outside Times Square and what they could do with it.
**swyx**
Beautiful. Amazing. I always thought that protein folding would be solved by them, but it would require a special machine.
**swyx**
Maybe the government would buy 5 of these things, and we could fold maybe 1 protein a day or 2 proteins a day.
**Alessio Fanelli**
And when AlphaFold came out and it was like, âYou can do it in Google Colab, on a GPU or desktop,â it was so mind-blowing. I forgot that protein folding was solved. I always thought that was inevitable, but the fact that it was solved and you could do it on your desktop just completely floored me. It changed everything.
**swyx**
Yeah. I don't even know what it is, but imagine ChatGPT came out, but instead it was like, âOh, you can just run it on your phone or locally on your own desktop.â That's the level of shock that came out.
**Alessio Fanelli**
And it gets down to this thing that humans are really bad at estimating problems that aren't human-made problems. Protein folding, we all thought, would require a huge amount of computeâa very challenging problem, the hardest problem in the world, right? It turns out that you can actually do it with, I think, around 10,000 GPU hours. You can train a good protein-folding model. It actually turned out to be barely an inconvenience.
**swyx**
Therefore, why not?
**Alessio Fanelli**
Oh. Therefore, protein folding was highly efficient based on experimental data. They took X-ray crystallography data. That's what DeepMind did: they took X-ray crystallography data. D. E. Shaw Research tried the first-principles method, and it was a nice head-to-head comparison. Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first-principles simulation by a very large margin.
**swyx**
And so why isn't Boltz, or whatever, inside of Cosmos? Why isn't there a tool that can run?
**Alessio Fanelli**
Oh, we have Boltz insideâwe have BoltzGen. Yeah, we have that inside of Cosmos.
**swyx**
Okay.
**Alessio Fanelli**
I mean, I think in the version that we have for people to just sign up and use, it's not in there. But you can imagine that you can just use Modal or Lambda or Tamarind or 310. There are all these companies that basically wrap a lot of these deep-learning protein-design tools or chemistry-design tools in an API. You can give that to Claude Code if you want. You can give it to Cosmos and be like, âHey, if you want to design a protein for X, use these tools.â
**swyx**
Your mechanism, it sounds likeâor one of the primary mechanisms that has been successfulâis to enumerate a whole bunch of possibilities and filter, right? How do you think about serendipity and out-of-distribution thinking and getting there? How far have you gotten, and what's left?
**Alessio Fanelli**
That's a great question. I think the short answer is that this is the domain of CBRN: chemical, biological, radiological, and nuclear weapons, or, I don't know, safety. This domain has been explored a lot in history by a lot of organizations.
I would say that there was a big question mark for us a few years ago: how much of this stuff is intellectually bottlenecked? How often are people like, âOh, wow, I want to cause harm, but I need to know some facts,â and could LLMs make that easier or go faster or anything like that?
I think the first set of answers in 2023 was basically no. You can go find the synthesis route for many dangerous compounds on Wikipedia. People know what the targets in the human body are that are targeted by most biological weapons. It's not really that much of a mystery. So I don't think there was a lot of new ground when LLMs first came about.
Then there was a lot of concern about laboratory protocols: could agents or LLMs reveal some tacit knowledge that maybe people couldn't find on Wikipedia? Maybe for making something, there's some technique that's required when you scale it up in size, or maybe there's some way to get around tracking lists by ordering different compounds.
That, I think, was really well testedânot by me, but by a few different labs. Some groups spun up and started making tests for this, and labs pay attention to it. I think it's really been put into the process where LLMs will shut down or be filtered in those scenarios, but I think that is actually an area where there is some risk.
I think this is something that people pay attention to for open-source models, and there's still some discussion there, but to a large extent, it's not really greatly accelerating in practice, or at least I haven't seen much evidence of it. Again, I think it comes down to the fact that it's not really available, but if you look hard enough, you can find most of the information you would need to get up to no good in the public domain already.
Alessio Fanelli
Yeah.
swyx
But I think now the next frontier is: can it somehow help you with real-time protocols and troubleshooting, more in the loop, and especially on the computational side of things? There are some scenarios that are now coming into focus that could be more dangerous or more intellectually bottlenecked, and so I think people are trying to pay attention to that.
To some extent, there was a first wave where we thought this could unlock a lot of stuff, and I don't think it came to pass. I think there's now an emerging second wave: there are some actually new scenarios that were just too far-fetched to consider 2 years ago that I think are now realistic. Some smart people are paying attention to it, but I don't think it's solved yet.
Alessio Fanelli
I don't know. It's very vague.
swyx
No, I mean, I guess one kind of differentiator is that there's a lot of talk about AI safety in the modern LLM and ASI space, and there are jokes about paperclip-maximizing robots or something. But the core threat here is more like a malicious actor using this as a tool to accelerate something dangerous.
The first-order hypothesis is that you basically already have to be an expert to effectively create a biological weapon or a chemical weapon, and a non-expert wouldn't know how to do this. An expert would already know how to do this.
Alessio Fanelli
Yeah. I think each of the categories in CBRN is a little different, but to a large extent, it's a lot of pushing material around. The classical example in nuclear is that it's a lot of centrifugation, a lot of ultracentrifugation, and a lot of high pressure or high RPMs.
You can maybe get smarter about how to set up the economy of scale to do that with an LLM, but to a large extent, you can call your friend in country X and they can tell you what the steps are. It's not that much of a secret; it's just a lot of moving material around, and I don't think it's meaningfully accelerated.
Now, that said, there are all kinds of dumb dual-use things. Maybe you want to call a company that makes centrifuges, and you want to make sure that they sell them to you and go through some KYC steps, and maybe an LLM can get you through the KYC faster. That's a dumb thing where, yes, email makes it so you can order centrifuges off the internet more easily. Is email a dual-use technology? Yeah, to some extent it is.
And so I think there are a lot of weird second-order things that we don't pay attention to in AI safety: does it make KYC easier? Does it make it easier for people to know where to order this from, what the expected price is, or what they should order first? All those simple logistical things are accelerated by AI, just as a consequence of AI being an accelerating technology.
Certainly, guys, there's some scary stuff, and I try not to think about it too much.
swyx
Yeah.
Alessio Fanelli
I don't know. I guess I don't want to get too political, but I do think that right now the United States government is maybe taking a slower, less intensive look at safety. But there are definitely people in other spaces than the U.S. government thinking about it hard.
swyx
And do you think this is something people need to spend more time on? I do get waves of angst about AI, and I'm sure many people living in San Francisco get a little bit of them too. Sometimes I think there isn't enough work being done on it, and then sometimes I think, âWow, I need to mellow out. We have lots of time to think about it.â
What is my opinion on it, then? I don't know. I think my opinion is not fully formed. Yeah, you and Sam have done a lot of thinking about funding science and the future of science. You've been vocal about the reproducibility crisis and other things. First question: why this focused research organization, or FRO? What does that get you that you don't get from academia or a big lab or whatever?
Alessio Fanelli
A nice network of people. Of course, I think Edison is going to do great, but I think it's a mystery what's going to happen. I don't think we've had as much friction there as you might expect.
But yeah, this is all stuff that Sam and I think about all the time: how do you balance stuff like this? How do you balance the economics? There are some venture-backed companies that are having cash salaries over $1,000,000.
Alessio Fanelli
And itâs insane to me.
swyx
Yeah.
Alessio Fanelli
That you would use all of your cash from your equity financing on these insane salaries. In terms of total spend on GPUs, that can still be a small fraction of your burn. So sometimes it kind of makes sense.
swyx
Yeah, yeah. Thatâs one way to think about it. This is a good lead-in to the fact that youâre automating science in some capacity. Where does that leave scientists?
Alessio Fanelli
I think this is Jevons paradox we can try here. Let me start with a contrast: if we automate taxicab drivers, thereâs not going to be an increase in people needing to go places. Maybe thereâll be somewhat of an increase, but thereâs a finite amount of time people will be spending in cars, so thereâs an upper limit. When you automate that, itâs a scarcity thing; youâre basically displacing jobs when you automate driving.
In science, I donât think thereâs a finite appetite or a finite capacity for science. I donât think science is a scarcity thing. Itâs not like there are 100 more discoveries left to be made and then weâll be done. If we can make science go much, much faster, there will be no decrease in demand. There will actually, I think, be an increase in demand that matches whatever amount of automation we have.
My vision for what a scientist would be in the future is that theyâll be agent wranglers or Cosmos wranglers. Theyâll be exploring 100 ideas simultaneously, or working with systems like ours to make 10x the discoveries, 100x the discoveries, because I think thereâs an unlimited amount of scientific discoveries to be made. Thereâs no scarcity state where weâll basically displace them all. Thatâs what I would tell a first-year PhD student: everythingâs going to be just fine.
Then, when it gets into the nuts and bolts, I do agree that this is going to be a really hard thing. If I am the CEO of a company that makes scienceâa pharma company, a materials science company, or an R&D arm at IBMâI might think, âWell, I could spend $1 million more on compute for the AI scientist, or I could hire 10 more people.â I might just choose to go with the AI scientist because, to a large extent, hiring people is hard, right? Hiring an AI scientist is probably a little bit easier.
swyx
Yeah.
Alessio Fanelli
So I think there could be some friction. Another thing is that science is, in some ways, closer to art, in the sense that a large number of people appreciate good science. If you get published in Nature, itâs not because itâs necessarily going to be world-changing. Of course, thatâs part of it, but itâs also because people say, âWow, this is really interesting science.â
Alessio Fanelli
Yeah.
swyx
swyx
Yeah.
Alessio Fanelli
The people who enjoy science are also scientists. I think itâs kind of hard to imagine a scenario where there arenât scientists as the consumers of science. If theyâre going to be consumers of science, theyâre also going to be some of the producers involved in the process itself, right? If that makes any sense.
swyx
Yeah, youâve touched on this. The question in my mind is: what does a scientist do, then?
Alessio Fanelli
Thereâs a great short story by Ted Chiang, I think from around 2003, At first, scientists were displaced, and they became interpreters of what the AI scientists were doing. They read the AI scientistsâ papers and translated them for popular science or something.
Then they couldnât read the papers anymore, so they were left behind. They had nothing to do and just sat around.
swyx
But the problem is that
Alessio Fanelli
Science is something you have to translate to make any impact. Science cannot exist by itself. I do agree that engineering can exist by itself. If you give some system a goal, like making me a material that I can make a space elevator out of, you could not participate at the beginning or in the middle of the process. You could just come in at the end and say, âOkay, follow this recipe.â
But scienceâwhatâs the origin of life? Is there water on other planets? Why is one catalyst better than another catalyst?âthat has to hit human eyes and human brains at some point. So I think a human has to be involved in the process.
swyx
I donât want to be contrarian, butâ
Alessio Fanelli
Yeah, be contrary.
swyx
Why does a human have to be involved?
Alessio Fanelli
Why does a human have to be involved? Well, a human has to be involved at least at some point to say, âYes, this is good science,â or, âThis is bad science.â
swyx
Okay, so it goes back to taste.
Alessio Fanelli
Yeah. But I donât know. Maybe youâre right. Maybe thereâs no point for humans. Maybe itâll be like Sora, the AI slop app. But I think in Sora there are still humans at the end clicking the videos or something.
swyx
Yeah. The Sora analogy brings up an interesting point. Is it possible that, due to the biases of AI science, if we really go all-in on science, there will still be a market for boutique human science? There are still people who want to paint things the old-fashioned way.
More to the point, does it become even more important to have a human actively doing their own exploration because there will be large blind spots and biases due to the modelsâthings youâll never be able to overcome because theyâre baked into the training data? Without a human, youâll always get stuck in a blind spot that youâll never be able to overcomeâ
Araceli Biosciences, which is a company in Oakland or Emeryville, does really cool stuff with automation. I think theyâre going to be testing this theory. If thatâs the bottleneck, weâll be able to see evidence of it because theyâre going to start doing really well.
It could be true.
Mm-hmm.
I still want to say that all of those, in my mind, are scoped in terms of R&D for pharma or biology. None of them are attempting to answer big, fundamental questions. Maybe there are different levels to think about. It seems like the focus of FutureHouse and Edison is much more toward R&D and sort of end-run science.
I have some background in fundamental physics. Is there any thought about how to take on dark-matter candidates?
I just think the data to really give us a complete story is not there yet.
You know what? Iâm sure everybody at every company is the biggest critic of their own product.
Yeah. We think Cosmos is great, but thereâs a very large amount of room for improvement.
With Cosmos, thereâs an open-access version for everybody.
Yeah.
Do you provide access to other labs through a less open version?
We have a version of Cosmos with bigger resources. It can run for longer and it uses GPUs. When it does data analysis, itâll have a GPU. We use that for things like machine-learning experiments. If you want to know whether itâs better to pretrain first on noisy data or not, for example, we have prerelease models that are coming out, and we try those.
So, yes, we do. We also have research partnerships with companies where we build something specific for them, and that is something we think about. But broadly, I would say Cosmos, the version thatâs on the website, is pretty close to the best we have internally.
Yeah. I have a question. You previously stated that you think language is the naturalâ
Language of chemistry?
The future of chemistry is language. Yeah, yeah. So I wonder: do you still believe that?
Good question. I would say yes, I still believe that. In that opinion article, my point was that, at the time when I wrote itâwhich I think was maybe 3 years ago, perhaps 2023âwe had models for predicting the solubility of compounds, data about very large populations, papers, and code. The only way to bridge all that information is natural language.
The argument was that whenever humans canât bridge informationâif I canât talk about my code or some idea to youâIâll invent words until I can get the point across. Humans are always innovating on language to make it represent all known observations. People innovate on language to represent whatever code pattern they have. Coming up with words to represent everything we know is the only shared activity weâve been doing for this long.
For that reason, I think natural language is the only possible way to connect all the different pieces of data we need in biology, medicine, or any domain for that matter.
I think there are some caveats to this. If Yann LeCun were here, he would make an argument about world models, vision, or embodiedness, right? There are arguments against natural language: maybe thereâs something more that it does. Itâs not the complete story, or maybe natural language imposes limitations that you cannot exceed because youâre stuck in this abstract space that was invented by humans, and you canât escape it until you can touch something.
Yeah. I mean, it is an abstraction, right? Scientists basically work exclusively in abstractions to some degree. I find that interesting because, as you said, most scientists, when they explain things, explain them through language, but many conversationsâmaybe mostâat some point result in people drawing diagrams or something. Chemistry, biochemistry largely, or medicinal chemistry, is often a language of graphs, right? Bonds are abstractions, yes, but theyâre pretty good abstractions for many cases.
Or geometry: think about a protein as the geometry of a protein. I think thatâs how a lot of scientists like to think about things. I find it interesting that youâre focusing primarily on language. Have you thought about a multimodal version of this, where, when it comes to a SMILES string, it doesnât just say, âOh, this is a SMILES string,â but, âThis is a graph; this is a representation of some higher abstract objectâ?
Youâre absolutely right. The problem with this Jacobâs ladder, or whatever you want to call it, is that, yes, you can call a molecule by its name; you can show the graph. Then if you go to a molecule like ferrocene, it doesnât really have bonds in part of it, and so youâre like, well, we need to draw it visually.
Then you go to a molecule like, I donât know, glycine betaine: thereâs this dihedral angle, and so itâs not actually this thing I drew; itâs actually an ensemble between this thing and this thing, right? Then you go to benzene, and youâre like, well, not only is it an ensemble of these different conformers, it actually has electron density. You canât really ignore the electron density in benzene; you need to treat it correctly.
Well, you canât actually represent the electron density that way. You have to look at the correlation of the electrons individually, right? Because you canât really model benzene with DFT, right, or a functional. You have to actually look at the electron correlation. Electron correlationâwell, you can model correlation, but actually, when these things are in a solution, they have relativistic effects because thereâs a whole bunch of stuff around. So you really have to have relativity in there.
Youâre like, well, youâve got the relativity and the electron correlation, you have the bonds, you have the conformers, but you really need to think about the cosmic radiation background because it does actually impact everything, and there is some energy there, right? Before you know it, youâve run out of compute or whatever resource youâre using to model this.
So I think you have to draw the line somewhere. Natural language, like I said, is something that humans have worked for a long time to make into the least abstractâor, whatâs the word? Itâs somewhere on the border: itâs still abstract enough that you donât need to know all these details, but itâs still granular enough, or concretized enough, that you actually can make use of it.
There may be some other representation. Multimodal might turn out to be video, or maybe thereâs some other fusion that you can make. I like natural language because we all work really hard to make it right at that boundary. I do agree sometimes ideas slip and they canât be expressed in language; you have to get out the whiteboard, or ideas slip and you have to wave your hands around. Maybe then you need that degree of freedom to communicate.
Just digging in on this a little bit more, famously quantum mechanics is indescribable, right? Thereâs an argument that you cannot understand quantum mechanics with words, or with our preconceived understanding of the physical world, because it doesnât behave like the macroscopic world, and so the only way to understand it is through mathematics, right? I largely see language as the joint key of science as well, but I wonder if thatâs not true for many domains, and quantum mechanics is just the one that hits you in the face.
I mean, I donât know. I think there are 7 principles of quantum mechanics, or 5 or something like this, that you can actually express pretty concisely in language. I agree that you need to actually look at the consequences of them; you need some mathematics. I donât know. This is a challenge. I think you could actually describe a lot of quantum mechanics in language.
Sure. Sure.
But I see your point. I guess Iâm a realist. When I talk to my kids, maybe Iâll say, âOkay, let me draw for you.â I donât make sure that everything in our house is described with natural language, so I agree with you there. I think maybe we can be a little flexible with natural language and include equations and SMILES strings in it, and I think we can get a little bit farther. Maybe thatâs okay.
But some people like optionality: âIt could be this or it could be that.â Iâm somebody who likes to take strong opinions and see how much farther they can get me. In my career, itâs actually been better for me to take strong opinions which, in the deepest part of my heart, I know may not be correct or may not be fully correct. But once you take these strong opinions, you can move many steps down the road.
For example, at FutureHouse, we took the opinion that scientific agents are the future, and that allows you to skip a lot of steps, because a lot of other people were like, âWe need to build a foundation model for X.â
Yeah.
It may not be a correct opinion. It may be more subtle or more complicated, but itâs allowed me to get very far. Iâll drop it someday and maybe find a new one. Yeah, not yet, though. Thatâs my main opinion on the matter.
The Ether Zero story on your blogâI find it hilarious and kind of awesome. You know, when I was a kid, I loved the genie/monkeyâs-paw concept: be careful what you wish for, because you just might get it. Can you just talk about that? It was just a really fun story.
Yeah. Ether Zero was a hell of a project because, conceptually, it was a very short project: âHey, people have made a lot of progress in verifiable rewards in math and in computer science and code. Letâs see if we can do it in chemistry.â
Chemistry is not a verifiable field, right? Of course, you can go test something in the lab, but then we had to think about all these ways to make chemistry verifiable. One of the ones we settled on was: make a molecule that has 3 nitrogens, 2 oxygens, 10 hydrogens, or something. We thought that was a pretty verifiable question.
But every time we would train a model, it would find some new, insanely weird trick to generate these molecules. Iâll tell you one of the examples. It would make these molecules, and we would do some checks to make sure it had the right bonds, the right number of electrons, the right number of atoms, and stuff like that. But it would solve the problem in any way possible, right? It would put all the nitrogens over here, put all the oxygens over hereâjust things that donât look good.
And so we started coming up with these rules: letâs check to make sure it followed these good practices or those good practices. We found ourselves in this opposite-of-the-Bitter-Lesson situationâI donât know, the boutique lessonâwhere you try to make everything custom.
But one of the things it kept doing was putting these nitrogens in a row. It would put 1 nitrogen, 2 nitrogens, 3 nitrogens all in a chain. If you have 3 nitrogens, itâs explosive; 2 nitrogens is bad, and 4 nitrogens you canât make. It kept making these 6-nitrogen compounds, and theyâre literally impossible.
Many of the people on the team were computer scientists, and one of them sent me a message one day: âThis is on the cover of Nature today, on Natureâs website.â
Somebody made a 6-nitrogen compound, and this was somebodyâs career: to deliver this compound, because this is the most unstable, insane compound you can make. Itâs some ridiculous setup, and the spectroscopy to prove that was very difficult. I donât know how they did it. It was an amazing accomplishment. Look, Andrew, itâs not actually impossible.
It was so funny to me that our model was sitting here spitting out these 6-nitrogen compounds in 2024 or 2025, and the paper just happened to come out that year that humankind had finally made a 6-nitrogen compound.
So do you think those were actually synthesizable, even under these extreme circumstances?
No. No. Our model was just reward hacking.
Okay.
The model was so creative in ways to reward-hack. Another one we did was make sure that when it proposed a reactionââMake this compound. Tell me how to make this compoundââall the reagents were purchasable. You could purchase them; they werenât made up.
The reason we came up with that was that originally, it would just take the end compound, remove 1 atom, and say, âHereâsâbuy this,â and then put the atom on. Itâs like, okay, well, I wish it were like that. The reagents had to be purchasable, and then we thought it might be hard if they were all purchasable because sometimes you actually order things custom or something. So we said, âJust make sure 1 is purchasable.â
The first thing it started doing was putting nitrogen in there, because nitrogen is purchasable and it had no participation in the reaction. I was like, oh my God. Okay, it has to be purchasable, and it has to participate in the reaction. Then it started putting in acid-base chemistry. It would just put an acid here. Acids are purchasable, and it would move 1 atom. We were like, okay, fine, it canât be that. Everything has to be purchasable and participate in the reaction.
Then we found ourselvesâIâm sitting there one day building this ridiculous catalog of purchasable compounds and a Bloom filter so it could go fast enough in our training loopâand Iâm like, why am I doing this? How did I get here?
How did I get here?
It was really funny because pretrainingâor training transformers on just data, just supervised training where you have the inputs and outputs directlyâis very nice and relaxing. Things are always robust; things go pretty smoothly. When you do these verifiable rewards, where you have to write a bulletproof verifier, it is really difficult.
We had so many models trained only to find out that they were hacking some other random thing in our setup. Itâs really hard, and I donât envy the frontier labs that have to do this at a very massive scale, because we had a lot of adventures in Ether Zero. You guys should read the blog post.
Definitely read the blog post. It was a great read.
GRPO. We did make some modifications to GRPO.
Yeah.
I actually used to know all the names of these modifications, but I think DAPO is one modification, and the clipping we did was special. We explored a lot of that stuff.
Yeah.
It was also one of these things where you think the hyperparameters are wrong, the algorithm is wrong, and then you find out itâs just because you had somehow sorted the reagents when you made your training data, but in your test data you hadnât sorted them alphabetically. The model was just barfing because its whole strategy was to exploit something in the way you sorted things.
So, yeah, we explored a lot of different methods. I learned a lot about chemistry and nomenclature, and I actually learned a lot about medicinal chemistry as wellâmore than I ever wanted to.
Awesome.
Yeah. Thanks, Andrew, again.
Yeah. Thank you very much for joining us.