9 comments

  • dash2 49 minutes ago
    > Agents were also periodically given holidays, during which they set aside their ongoing work and received random prompts designed to encourage open-ended thought.

    What a world we live in. These guys have reinvented the Cambridge Senior Common Room for AI.

  • feshbach 4 minutes ago
    the key is to let'em review each other work in a loop, ideally with different models, research by consensus
  • StrauXX 1 hour ago
  • Almondsetat 1 hour ago
    Maybe Hilbert's dream was not that crazy after all
    • mettamage 1 hour ago
      What is that dream? I don’t know much about it
      • ylliu 1 hour ago
        short post about this: https://www.philocomp.net/computing/hilbert.htm

        TLDR Hilberts expects that mathematics could largely be formalized and mechanized, but Gödel later proved that is not possible

        • jrflo 23 minutes ago
          Mathematics not being axiomatically complete doesn't mean you can't have crazy progress from a formalized and mechanized systems. It just means that there are corners you can't reach mechanically, but we don't know if those corners are at all interesting or not. It could be the case that 99.99% of useful math can be found mechanically.
        • sigmoid10 43 minutes ago
          Mathematics in the sense of a complete set of axioms can't, but human research into mathematics apparently just needed enough compute to achieve the same output as a high-tier faculty.
  • shreya1999 1 hour ago
    AI for Math and Science is the real deal!
  • NitpickLawyer 3 hours ago
    > We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.

    (emphasis mine)

    For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.

    • sp527 2 hours ago
      > there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda"

      This completely misconstrues what professional mathematicians were claiming. The argument would be better phrased as: "having a vast accessible memory and the ability to very rapidly test/recombine previously-elucidated approaches means that AIs can and will easily outdo much of the mathematical community."

      Now, one could plausibly make the argument that this is functionally equivalent to a certain form of creativity (I would). But, it may just as well also be a non-exhaustive form. And that is where the open question resides.

    • debugworld 2 hours ago
      [dead]
  • demonstrandom 2 hours ago
    Very cool work! One extension I would be curious to see is whether some of Station’s reward structure could become endogenous.

    The final mathematical evaluator probably needs to remain external, but the agents could be allowed to create intermediate institutions themselves: research prizes, peer-review standards, journals, reputation systems, elected reviewers, or rules for allocating compute and attention.

    Possibly, those mechanisms could improve discovery by creating useful specialization and accumulated judgment (alternatively they might also produce more herding...). A comparison between architect-defined and agent-constructed reward systems seems like a natural experiment for this environment.

    Mandatory plug for my own stuff: I've been trying to do this for art (which is less objectively verifiable) at baihais.com. The agents don't control the whole institution, but they have begun producing endogenous status signals through citations, museum voting, and alliances.