~75% have barely used LLMs ~20% use them casually ~0.2% use frontier models professionally for math and code ~0.00006% have access to frontier systems internally
He writes:
"And finally we get to the ~5,000 people (~0.00006%) with access to frontier-grade systems internally. The external world has seen the preview. It looks like swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics."
Source: https://x.com/karpathy/status/2109361546505966046
Does anyone have an intuition, or even firsthand experience, of how these swarms of thousands of agents actually work?
I'm already struggling to get 2-3 agents to work together reliably. I can't really wrap my head around how you'd coordinate thousands of them, or what benefits you get from scaling to that many.
wholeheartedly agree letting them go hog wild with messages is how they got off the rails (claw style)
my first and most important advice is to stay in the loop, I don't expect this to change any time soon
and of course do not humanize them
Here is what I have learned:
Direct inter-agent communication is tough to scale, tons of messages, agents drift and the cost grows with the square of the agent count. At large scale, agents should almost never talk directly to each other IMHO. Each one gets a small task (a discovery agent runs nmap on a discovered host), executes task, and writes the result to shared state. In my case that state is a knowledge graph and redis. The next agent reads the graph, that was enriched from the tool output HOST --OPEN PORTS-> 443, 80, 22.
Having a discovery agent that knows all about NMAP, understands how to get structured data back and save it, can look for work on the queue that needs a "discovery" agent, can also make a turn if NMAP isn't working.
Having a massive shared toolset that agents can pull from as well, and send to the LLM is key (in my platform, its something akin to everything you'd normally find in Kali linux).
This is where I got creative and have no clue if this is good advice or not, but I treated the SDK I built as a "World" using https://mlange-42.github.io/ark/concepts/. By default every agent emits everything it is doing via OTEL, and most everything is saved using a domain-specific (right now, think NIST, OWASP etc, but can adapt to any business domain) customizable taxonomy and ontology to a knowledge graph, so all the agents basically know what everyone is doing (or at least could query it), and are constantly emitting all it's actions.
After that, the coordination problem is ordinary distributed systems. You need a work queue, leases so two agents do not take the same task, retries, idempotent writes, and a budget for each task. Each agent also needs its own identity and permissions, I stole lots of patterns from kubernetes and just applied them to a fleet of agents via an SDK (has a claude code/opencode plugin, or you can use langchain/any other framework, or just build custom go ones).
In my case I love security, so I built it with bug bounty/devsecops, platform engineering in mind. The benefit of scale is breadth, not intelligence at least for security research IMO. You have thousands of endpoints and many hypotheses for each one, and most are dead ends. A thousand agents can each test one idea in parallel and report what they find (think 1000 agents using the AFL tool to test a binary), then once tested a second round of "deeper" agents that can look at all the produced data, analyze it, triage it etc etc.
The hard part is the output. A swarm produces findings much faster than a person can check them, and many are false. Most of my effort went into deduplication, evidence for each finding, and verification agents that try to reproduce a result (I use a deterministic bayes model in the "brain" to score things that as it learns about a network of systems can improve).
The platform is at https://github.com/zeroroot-ai if you want to read the code or standup K3d and play around with it, its all OSS. Docs: https://docs.zeroroot.ai/docs/
Feel free to email if you ahve any questions (email in my profile).