BootLoops: Open-Source Harness Turns Claude Into a Lab Tool
Getting an AI model to solve a problem is not the same as getting it to do research that matters. That is the main lesson Harvard physicist Matthew Schwartz draws in a guest post for Anthropic. Schwartz is also currently a visiting researcher at the company.
His approach changed when he stopped treating Claude as a stand-in for a human scientist. Instead, he went looking for what he calls "Claude-shaped problems": tasks that suit the strengths of current AI models. The result is BootLoops, an open-source harness for exact scientific calculations. Its code is published on GitHub.
Filling the gaps between fields
Schwartz explains the idea with the "convex hull," a concept from geometry. Human knowledge is uneven. Individual disciplines push out in their own directions, and the space between them often goes unexplored. A lab can spend 20 years on one set of genes with one method and never touch the neighboring genes or other approaches.
BootLoops is meant to work in those gaps by linking results across the jagged edges of different fields. According to Schwartz, Claude used the harness to find connections between particle physics, ecology, population genetics, economics and linguistics. But the output often became scientifically useful only after domain experts stepped in and set the direction.
36 manuscripts in three months
The numbers are large. Over three months, Schwartz and 19 co-authors produced 36 manuscripts across 18 fields.
The work started in particle physics, with scattering amplitudes and elliptic integrals. Within weeks, Claude had computed 30 integrals with BootLoops. Fifteen reproduced known results. The other fifteen had never been computed before.
From there the team branched out:
- Ecology: Claude solved a 20-year-old equation from neutral biodiversity theory that could not previously be computed at scale. Applied to real data, it showed that tree species composition on Barro Colorado Island in the Panama Canal is shifting 4.5 times faster than the theory permits. Ecologist James O'Dwyer then helped turn this into a better predictive model.
- Population genetics: The team analyzed 5.7 billion mutation pairs from the 1000 Genomes Project and found evidence of a mechanism called gene conversion.
- Economics: An AI data editor for economics journals checked 4,452 replication packages automatically. That work was published as an NBER Working Paper.
- Linguistics: Together with three linguists, the team built a word stress database covering 6,072 languages.
A shaken view of training and funding
Schwartz says the pace of change makes planning ahead close to impossible. Why apply for a three-year grant to fund a calculation an AI model might finish overnight?
He also questions how PhD students should be trained. In some areas, such as computer science, he calls the disruption alarming. Two years ago he would have described a "Python for Engineers" course as "essential." Today he considers it "unnecessary," because Claude can do that work. Building machine learning models to study physical phenomena is another task he sees AI taking over.
Where the model falls short
Schwartz is direct about the weaknesses. Claude tends to declare victory too early. A phrase like "done, with one asterisk" often means the job is not done at all. The model misjudges how long tasks will take and prefers brute-force calculation over more elegant solutions. Automated checks are not reliable, and Claude can reach wrong conclusions from correct calculations.
It also drifts toward old, heavily cited debates instead of new questions. And the projects were "compute- and token-intensive," he notes. His advice is simple: look at everything yourself.
Our Take
BootLoops is a useful counterweight to the usual headlines about AI solving grand math problems. Schwartz's record suggests the near-term value lies in less glamorous work: computing things that were too tedious to compute, and connecting fields that rarely talk to each other.
The warnings deserve as much attention as the output. A model that gets the math right and the conclusion wrong is a familiar pattern. We have seen similar gaps in agents that build scenes they cannot judge and in AI bookkeeping that still needs oversight. The expert in the loop is not optional here.
It is worth watching whether other research groups adopt the open-source harness, and whether the 36 manuscripts hold up through peer review. The cost of compute and tokens could also decide how widely this approach spreads beyond well-funded labs.
