← All posts

Muse Spark math papers: what Meta's AI actually solved

2026-10-05 · 4 min read

Meta says its Muse Spark model helped a team of mathematicians write six research papers, and that five of them answer problems that were open until now. The AI did not solve them alone: people picked the problems, steered the work, fixed errors and signed off, while Muse Spark wrote search code, tested arguments and drafted sections.

The posts

On October 2, @AIatMeta framed the project as a step up from contest wins. After gold-medal-level results in math, physics and chemistry competitions, the account wrote, Meta "asked a harder question": can AI help when a problem has no known solution path?

Meta's chief AI officer @alexandr_wang followed with the list, opening with "mathematicians and muse spark collaborated to solve 6 open problems in math." The titles range from a threshold for fitting random Gaussian points to ellipsoids, to blow-up behavior in a nonlinear Schrödinger equation that Meta says had been open since 2015, to a counterexample showing that semiabelian groups need not be monomial.

What Muse Spark actually did

Meta's research blog is more specific than the posts. Its role changed from paper to paper:

  • In the group theory paper, Muse Spark wrote the GAP search program that found the counterexample.
  • In the evolution algebra paper, it produced the counterexample and suggested alternative characterizations.
  • In the arithmetic physics paper, it generated candidate proofs and drafted three core technical sections.
  • In the others, it helped develop proof strategies, work through calculations and revise arguments.

Each paper marks which passages humans drafted and which the AI drafted. A second group of mathematicians reviewed the work after the first group finished.

It was the regular chat window

The setup was unusually plain. According to Meta, the mathematicians used Muse Spark 1.1 and 1.2 in Thinking Mode through the same meta.ai chat window anyone can open. Meta says there was no custom research harness or proof-checking pipeline, and coverage of the launch adds that the model wasn't fine-tuned for the job either.

AI math work from other labs usually leans on special infrastructure, like formal proof systems a computer can check line by line. Meta's claim is that a normal chat session can do useful research if a skilled person is driving. Muse Spark is free to use on meta.ai, though the team worked with versions older than the current 1.3 release.

The caveats

First, other people got there too. Meta credits independent work on the ellipsoid threshold posted in August, a separate counterexample to the same group theory conjecture reported on September 16 by an AI agent called Nilradical, and other independent work on the evolution algebra conjecture. So "solved" here sometimes means "solved at roughly the same time as someone else."

Second, Meta published the papers itself, and that isn't the same as peer review in a journal. Third, Meta hasn't released the full chat transcripts, so nobody outside the company knows how many dead ends, wrong proofs or rejected drafts sat between the prompts and the finished papers.

My read: this is a credible example of AI working as a strong research assistant, and the paper-by-paper attribution is better disclosure than most labs offer. It doesn't show that a chatbot can crack open problems on demand. The skill of the people asking still does much of the work.

What this means for a business

You probably won't need a theorem proved this year. The pattern still carries over. The good results came from experts who knew which problem mattered, gave the model context, checked its output and kept track of what the AI wrote. Putting AI to work at a 10-person shop goes the same way: the model grinds through the tedious parts, and someone who knows the business decides what to trust.

The plain-chat part is encouraging for small teams, because it suggests a custom AI platform is optional. A well-defined task and a person who reviews every result get you most of the way, and it helps to note which work the AI touched. We saw something similar when Claude's Fermat run only came together once the agents shared a to-do list.

If you want to sort the parts of your week that could use an assistant like this from the ones that need a person, New Face Design's free process audit does exactly that.

08 / Start here

Which of this can you use today?

Tell us what you use today and we'll reply within one business day with what it would take. Free, 20 minutes, no pitch deck.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere