HEP-TH and AI

I’ve been talking to people and seeing a lot of discussion about the effects of AI agents on math research, but haven’t found much serious discussion of what is happening in hep-th research. I’m curious to hear from those involved about exactly what is happening as researchers try and exploit these new tools.

Back in February I tried to generate some data about trends in hep-th submissions, initially got this wrong by using “recently modified”, rather than “initial submission” dates. I’ve started playing with some AI agents, and one thing they’re good for is generating python scripts to gather and plot this kind of data. Here’s what hep-th submissions (not including cross-listings) look like for the past three months, compared to earlier years.

There’s a very clear pattern: the rate of hep-th submissions is now about 1/3 higher this year than in previous years.

I haven’t tried to count submissions acknowledging AI use, and from a quick look at the submissions, there are relatively few that explicitly advertise dependence on AI generated results.

One thing that is showing up is that, in hep-th as elsewhere, AI is being used to scam the public. A new press release from King’s College London announces that String theory finally testable with power of AI, which is based on this hep-th preprint, recently published in Phys. Rev. D. The press release includes:

Overturning decades of popular conjecture that theoretical string theory cannot be tested, scientists have robustly proved that it can be for the first time…

Author Dr David Marsh, Ernest Rutherford Fellow at King’s College London, said “The popular understanding that you can’t experimentally test string theory is no longer true. While previous studies have tested a few different models of string theory, our computational approach improves on this with rigorous testing across millions of different models.

“Rigorous falsification is the core of scientific experiment, and by finally applying this to string theory we finally have the tools to test this vaunted ‘theory of everything’. Finally touching experimentation here is an exciting step forward for theoretical particle physics.”

So, it’s clear that AI agents can produce hep-th papers designed to be promoted as outrageous bullshit, but I’m interested to hear about the opposite end of the credibility spectrum. What are examples of serious advances in the subject that are coming from AI tools?

This entry was posted in AI, This Week's Hype. Bookmark the permalink.

21 Responses to HEP-TH and AI

  1. Zohar says:

    It is extremely useful for numerics, and it is substantially accelerating a project that I am working on, where heavy numerics is required. I have not seen it produce deep thoughts or interesting ideas, yet. In fact when there is some conceptual confusion or some deep issue, it is better to avoid talking to the AI since in 99% of the cases it would waste a colossal amount of your time in vain.

  2. Architrino says:

    I think it will be challenging for Ai to help hep-th directly. My view is based on the thought that what we observe and theorize about at high energies, high frequencies, and small distances — could very well be many orders of magnitude away from the real action. Or in other words, everything we know today could be bulk properties of matter and ‘spacetime’ which operate under very different rules at their scales of action, yet lead to what we observe. So, unlike mathematics, hep-th may find itself in a ‘you can’t get there from here’ situation with Ai. If all the training data is steeped in current state of the art, there may be very few or difficult to discern patterns that hint at the underlying behaviour in a way that Ai could recognize.

  3. Paul says:

    I would plot Jan to May 2026 by month to see if there is a steady progression. Sudden jump at start of year might be suspicious.

  4. Peter Woit says:

    Zohar and Architrino,
    Thanks for your comments, they’re quite interesting.

    Paul,
    The plots I gave were produced by asking chatgpt to do it. It quickly produces an appropriate python script, so anyone who wants to can pretty easily get a script that will generate plots for whatever ways you want to look at these submission numbers.

  5. Andrei says:

    You know https://vibemathed.com/

    It has a section “mathematical physics”

  6. Scott Caveny says:

    Here are a few links concerning use of ML for research in theoretical physics:
    1. Matthew Schwartz March 2026 review of Claude: https://www.anthropic.com/research/vibe-physics
    2. Living review of ML for High Energy Physics (June 2026): https://iml-wg.github.io/HEPML-LivingReview/#ml-for-theory
    3. NSF AI Institute for Artificial Intelligence and Fundamental Interactions: https://iaifi.org/research.html

  7. 4gravitons says:

    Paul, someone already did that kind of analysis here: https://x.com/nblqbl/status/2076635168727577047

    It looks like a takeoff between April and May, still growing as of those plots. It struck me as newsworthy enough to pitch to a few places at the time, no luck though.

  8. A Slopocalypse Skeptic says:

    Paul Ginsparg recently gave a Cornell colloquium on “The Rise of Slop”. I’m not sure if it was recorded, but he discusses the “slopocalypse” starting around 11:18 in the video here (at Lance Dixon’s 65th birthday conference):

    https://www.cs.cornell.edu/~ginsparg/pg_lancefest.mp4

    Note that for the most part he talks about arxiv as a whole, not hep-th specifically, and in fact the problem seems to be worse in CS than in physics.

  9. Greg says:

    I found this blog by a Hungarian physicist (Balasz Pozsgay) to be very informative (even to a mathematician) both as far as norms and as far as capabilities are concerned: https://aiforintegrability.substack.com/

  10. Greg says:

    Here is a paper on the arxiv today (https://arxiv.org/pdf/2608.15505) resolving the remaining cases of Milnor’s Ricci-pi_1 conjecture. The conjecture was that complete Riemannian manifolds with non-negative Ricci should have finitely generated fundamental group. This had been proved in dimension 3 a decade or so ago and counterexamples in dimensions greater than 6 were constructed a couple years ago by Brue, Naber and Semola (this was an annals paper). They later improved this to dimension 6, which remained open (and believed to be hard, though I am definitely not an expert). The preprint claims to construct counterexamples in all dimensions greater than 3. It is a long paper (62 pages) disclosing significant ai use for brainstorming, objection generation, and computation checks (along with more mundane stuff like literature sourcing).

    I expect this type of `hybrid’ paper to become common in math in the next weeks and months, and am curious whether or not theoretical physics will follow a similar route.

  11. Peter Woit says:

    Greg,
    Thanks, that’s interesting. But, here I really want to generate discussion about what is happening in hep-th, where things seem to me to be very different than in pure math research.

    All,
    Please, here just stick to comments about hep-th. There will be other opportunities to discuss the evolving pure math story.

  12. Emil Martinec says:

    I’ve been working with Claude since our university acquired a site license a few weeks ago. It is tremendously useful for numerics, something I have been hesitant to enter into because to do it well requires more of an investment than I was willing to make. It is well-versed in all sorts of numerical methods. But it needs a firm guiding hand because it has zero physical intuition. What seems to work best is to describe the problem you wish to attack, ask it to read relevant background material and describe the task in its own words, what it thinks it should do, and to make a research plan. Then discuss and revise the plan so that it doesn’t launch off in some random direction that you didn’t intend. Tackle the problem in discrete stages with clear milestones and making course corrections along the way so it doesn’t spin its wheels. On the symbolic manipulation side, it has also been useful; I say half-jokingly that it makes mistakes an order of magnitude faster than I can. Which means you can explore all sorts of possibilties much faster, and focus on the most promising direction. If you give it a textbook problem that is commonly seen it knows exactly what to do and will do it correctly, but is again lost on long chains of inference without being given some guidance. I see AI as a tremendously useful tool already for the problem I am currently working on, but it is not yet able to have intuition or to be creative, though it has surprised me once or twice in what it was able to come up with. Time will tell whether these systems gain those capabilities. As for the question of what it will do to hep-th, yes the subject will see an increase in the number of uninspired and/or derivative and/or possibly wrong-headed work, because AI will be able to help generate it faster than humans currently can (my comment above about being able to make mistakes faster). Hopefully AI will be useful in helping sort through the mountain of useless results that is sure to come ;^)

  13. David Marsh says:

    We didn’t use an AI agent to write the paper. The “AI” is really ML. We used normalizing flows as a way to speed up Bayesian inference compared to “classical” MCMC.

    Also you and Sabine both totally fell for my bait. Like shooting fish….

  14. Peter Woit says:

    David Marsh,
    Thanks! This makes clearer where string theory has ended up: a massive trolling operation….

  15. David:

    Now I’m confused. I can’t follow the Arxiv paper at all–too much technical detail–but I did get that it was doing Bayesian inference. But then I didn’t understand why the press release said that string theory is “finally testable.” At best, what would be testable here is the specific model being fit here to data, not the string theory model itself, right?

    And what does it mean for an article or a press release to be “bait” in this context?

  16. Chethan Krishnan says:

    I use AI to generate first drafts of sections now. It is a massive time-saver, especially in review-heavy content. But I also think it is largely uninspiring work. In places that matter, you have to edit it tons, but it still saves some time.

    The place where I think AI needs a real hat tip, is when you are trying to learn something new that you are not an expert at. I feel that it has really broken down the barrier-to-entry into many subfields, for those who are not in the specific clique. Sure, there is a diminishing returns aspect to AI as you get to the frontier, at least in hep-th related matters, but by then I can stand on my own, so it’s fine. Previously, if an idea took me somewhere that was way outside my comfort zone, I would often convince myself that the idea wasn’t worth pursuing. I would drop the joke, rather than find the pen-and-pad, to paraphrase Mitch Hedberg. I do it less now.

    This is also why in my acknowledgements, I thank AI **along** with my colleagues and friends that I have discussed things with: you learn from them. Generating text, at least in the way I am using it, doesn’t get to the heart of what is useful for me, about AI. It is just faster word-processing.

    As for coding, my students used AI to speed up runtime of some code they had written, which was useful for a paper we put out in March. I am currently using Claude to generate some python code for getting hints about what I hope will eventually turn into a calculaiton, but the jury is still out.

  17. Sabine says:

    @David Marsch

    I didn’t so much as comment on your use of the term “AI”, so wtf are you even talking about other than admitting that you deliberately lie to the public. People like you are an embarrassment for the entire discipline.

  18. jsm says:

    One absolute rule: no one should paste anything written by AI into a manuscript. Make sure it first passes through your brain (and fingers).

  19. Armando says:

    Matthew Schwartz’s article is the most relevant from the ones I know and it should really get more comment here for a few reasons:
    In the space of three months the AI went from not being able to solve the problem, to solving it more than 20 times faster than a smart graduate student (Schwartz’s estimate not mine)
    The AI lied a lot during the process and had to also be hand held. Graduate students in principle should lie less and will also need handheld
    The AI was able to write a paper that is actually decent to write and not the usual myriad of sentences that seem to make sense but actually don’t.

    If you haven’t take a serious look into Matthew Schwartz’s blog post and the produced paper.

  20. Eleni Petrakou says:

    This comments’ thread has been an inspiration,
    https://thephword.substack.com/p/how-to-not-write-scientific-press

Leave a Reply

Informed comments relevant to the posting are very welcome and strongly encouraged. Comments that just add noise and/or hostility are not. Off-topic comments better be interesting... In addition, remember that this is not a general physics discussion board, or a place for people to promote their favorite ideas about fundamental physics. Your email address will not be published. Required fields are marked *