RSA-260 Fell to an Agent Fleet. The Human Sent 3,328 Messages.
If you plan to point agents at a job that runs for weeks, read the RSA-260 factorization first. It is the most detailed public record of such a run. The work was cut into ten-minute units that any machine could pick up. Arithmetic set the checkpoint intervals. A verifier ran beside production, and one person spent most of his effort telling agents what to stop doing. This essay reads the run for those parts and gives you the commands, parameters and formulas to copy into your own long job.
At 01:48:57 UTC on September 3, a job on a GPU cluster printed two 130-digit primes. Their product is RSA-260, the largest number from the RSA Factoring Challenge that anyone has publicly split. Today Cognition published the full writeup. The first prompt went in on August 13. Nine hours later, an agent had a working GPU lattice siever. Three weeks after that, a fleet of agent sessions had spent about 4,900 GPU-days to run the general number field sieve to completion.
The steering log shows in unusual detail what kept agents productive on one computation for three weeks. That log, more than the cryptography, is the subject of this essay.
What the run computed
RSA-260 is a 260-digit semiprime that RSA Laboratories published in 1991 as a public challenge. The previous record, RSA-250, fell in February 2020 to an academic team that ran CPU clusters. The tool for numbers this size is the general number field sieve (GNFS), and its open-source reference is CADO-NFS. Cognition's run used CADO-NFS as the base and borrowed parts of msieve for polynomial selection. The agents replaced almost every compute-heavy stage with a GPU version that they wrote. Per the writeup, only four pieces stayed on CPU: the cado-nfs.py driver, polyselect_ropt, makefb, and dup1.
GNFS runs in five phases, and each phase has a different shape. Sieving splits into many independent jobs. Linear algebra is one long chain of steps.
- Polynomial selection searches for a pair of polynomials that make the later sieve cheap. The run generated 83,557,723 candidate polynomials, root-optimized 18,333, trial-sieved 22 and kept a degree 6 polynomial. This phase cost 643 GPU-days between August 18 and August 21.
- Sieving collects "relations". A relation is a pair of small integers for which two related values both factor into small primes. The work splits by a parameter called the special q. Cognition sieved q from 1.0e9 to about 3.91e10 in 632,249 workunits of 60,000 q each. Each workunit was sized to take about ten minutes. This phase cost 3,813 GPU-days and 189.5 wall-clock hours, from August 22 to August 30.
- Filtering removes duplicates and useless relations, then builds a matrix. The sieve produced 13,849,985,589 raw relations. Of these, 8,298,749,059 were unique (40.1% duplicates). The run also had 3,991,449 free relations.
- Linear algebra finds dependencies in a sparse matrix over GF(2) with block Wiedemann. The matrix was 656,182,601 by 656,182,189 with 98,431,741,898 nonzeros. This phase cost 467 GPU-days.
- Square root turns a dependency into an actual factor. The final numbers appear in this phase.
Add the three measured phases and you get 4,923 GPU-days. Cognition rounds this to 4,900, or 13.5 GPU-years. The writeup prices GPUs at $3.50 per GPU-hour. So 4,900 × 24 × $3.50 is $411,600, the "about $400k at current market prices" in the post. The planner below lets you move the two inputs that sized the expensive phase.
The sieve must produce enough unique relations to build the matrix. Duplicates are wasted work. So the duplicate rate sets how many workunits you pay for, and the GPU pool sets how long you wait. The defaults are the RSA-260 run: 40.08% duplicates, the exact share behind the rounded 40.1%.
The planner fixes per-workunit yield at the run's measured 13,849,985,589 / 632,249 = 21,906 raw relations. Sieving GPU-days scale with workunits from the measured 3,813. Polynomial selection (643) and linear algebra (467) stay at their measured values. Cost uses the unrounded 4,923 GPU-days. A real run with a different unique count would also change the matrix. The average of 483 GPUs comes from 3,813 GPU-days over 189.5 hours.
Nine hours to a GPU siever
The project started with a narrow request. Eric Lu, who wrote the post, quotes the opening prompt: "CADO-NFS is FOSS software for performing GNFS. I'd like you to develop a fast GPU lattice siever... You have Modal access keys available that I authorize you to use to spin up a single GPU box for performance testing." The agent worked for two hours, then iterated overnight for seven more. The result was glas, a drop-in replacement for las, the CPU lattice siever in CADO.
Sieving is the right place for an agent to start because the task has a hard oracle. A siever either produces valid relations at a measurable rate or it does not. To check any candidate change, run a workunit and count the output. "The intersection of GPU kernel-writing experts and number field theory experts is quite small," Lu notes. A fast, cheap verifier lets an agent substitute iteration for that rare combination of expertise.
Optimization then continued for weeks. The writeup reports per-workunit times that fell from 585 to 486 seconds on GB200 nodes. They fell from 586 to 504 on GB300 and from 628 to 541 on B200. That is roughly a 14 to 17 percent speedup on hardware that was already running. A team usually does this sort of tuning once and then stops.
Here the tuning ran in parallel with the production run. In Lu's words, the sieve is "embarrassingly parallel over billions of small work units." A faster siever could go in between workunits without stopping anything.
Build that oracle before you rent a GPU. CADO-NFS ships one in its README: a 60-digit test number that runs every phase end to end on one machine.
./cado-nfs.py 90377629292003121684002147101760858109247336549001090677693 -t 2
From the CADO-NFS README. -t 2 sets the thread count. Not run here.
Give the agent that command as its definition of done for any change to a stage. Add a rate measurement on one real workunit. With a task written this way, the agent cannot declare victory on a faster kernel that returns wrong relations:
Optimize glas, the GPU lattice siever.
After every change:
1. Run the c60 smoke test. The two factors must come out.
2. Run one RSA-260 workunit (q range 1.0e9 to 1.0e9+60000) and report raw relations per second.
Keep a change only if both hold and the rate went up. Log each kept change with its rate.
Our pattern, built from the opening prompt quoted in Cognition's writeup. The workunit range is the first 60,000 special-q of the run.
Sieving scales out, linear algebra does not
The two expensive phases look nothing alike to an operator. Sieving is 632,249 independent jobs. Any GPU of any type can take any workunit, and a lost workunit costs ten minutes. The job queue absorbs mixed hardware without complaint. Divide 3,813 GPU-days by 189.5 hours and the run averaged about 483 GPUs at once.
CADO-NFS runs this as a server that hands out workunits to clients that come and go. Here is the two-machine example from the README:
machine1$ ./cado-nfs.py --server 90377629292003121684002147101760858109247336549001090677693 server.port=4242 server.ssl=no server.whitelist=machine2
machine2$ ./cado-nfs-client.py --server=http://machine1:4242 --bindir=...
From the CADO-NFS README. The example disables SSL for a trusted network. On any other network, keep SSL on and pass the server's certificate to the client, as the README describes. Not run here.
Workunit size and the number of losses that the server tolerates are both parameters. The example parameter file sets them for a small test factorization:
tasks.sieve.qrange = 2000
tasks.wutimeout = 120
tasks.maxtimedout = 100
From CADO-NFS scripts/cadofactor/parameters. These values are the test defaults in the file. RSA-260 used other values.
The comments in that file carry the operational advice. wutimeout cancels and resubmits a workunit whose client went silent. So it must be longer than your slowest honest workunit. maxtimedout aborts the whole factorization after that many timeouts. The file says the default "is probably too small" for machines that stop and restart often. On a preemptible GPU pool, raise it before launch.
RSA-260 set its q range to 60,000 per workunit, so each workunit ran about ten minutes. For your own job, time one workunit at your target q and scale from there.
The driver is restartable if you name its work directory. The README shows how:
./cado-nfs.py 90377629292003121684002147101760858109247336549001090677693 workdir=/tmp/myfacto
./cado-nfs.py /tmp/myfacto/c60.parameters_snapshot.0
From the CADO-NFS README. The second command resumes an interrupted run from its parameters snapshot. Not run here.
Linear algebra is one computation. Block Wiedemann runs a Krylov sequence. This is a long chain of sparse matrix-times-vector products, where step n needs step n minus 1. Cognition ran 2,564,096 iterations per sequence and two sequences, each 256 wide. The MPI grids ranged from 4×2 to 16×4 nodes, and the phase ran for about 48 hours. On shared capacity, the post says, the job "was fatally preempted frequently and we had to develop a somewhat complex placement script to constantly fit the best shape possible to the available compute."
Checkpoints limit the damage. The run wrote a checkpoint every 8,192 iterations and kept every 32,768th for the later solution step. A preemption kills the job, and the work since the last checkpoint must run again. Sparse checkpoints mean a lot of redone work. Frequent checkpoints make the writes themselves the overhead.
CADO's own consistency check also fired about 14 hours into the Krylov phase, and the team re-ran that interval. In parallel, CADO's bwccheck verified every pair of checkpoints. That check is the only way the run could catch a silent corruption.
In CADO's block Wiedemann driver, both numbers are parameters. The bwc README gives this MPI invocation:
./build/xxxxxxx/linalg/bwc/bwc.pl :complete thr=3x4 mpi=2x2 matrix=/localdisk/tmp/nfs/c156.sparse.bin interval=4096 m=64 n=64 wdir=/tmp/bwc.wdir
From the bwc README, for a 156-digit example. Not run here.
The same README documents how to keep only the checkpoints that the later mksol step needs. Give krylov a checkpoint_precious equal to the mksol interval. With the run's numbers:
krylov: interval=8192 keep_rolling_checkpoints=4 checkpoint_precious=32768
mksol: interval=32768
Pattern from the documented combination in the bwc README. 8,192 and 32,768 are the values of the run, per the writeup. We chose keep_rolling_checkpoints=4. The README calls rolling checkpoints with skip_online_checks=1 dangerous: it "leaves open the possibility of a failure which stands no chance to be eventually detected."
One Krylov sequence, 2,564,096 iterations, drawn as a tape. Blue ticks are checkpoints. Red markers are preemptions. The red segment behind each one is work since the last checkpoint that must run again. Pick an interval and a preemption count, and the ledger prices the damage.
Rate is 2,564,096 iterations over 48 hours, about 14.8 per second. Interval 8,192 is the real setting of the run. The preemption count, their positions (fixed pseudo-random) and the seconds per checkpoint are illustrative. The writeup says preemptions were frequent, but it does not count them or time the writes.
Move the interval and the trade appears quickly. With no checkpoints, a single late preemption throws away most of two days. At 131,072 iterations, each preemption costs on average an hour or more of redone work. At 2,048 the redo nearly vanishes, but you write over a thousand checkpoints. If each write is slow, the writes then cost more than the preemptions they protect against.
The 8,192 of the run is a middle setting. With writes of a few tens of seconds, it keeps both the redo and the write bill small. Anyone who has run long MPI jobs knows this trade. In this run, an agent session operated the loop, including the placement script that reshaped the job to fit whatever capacity was free.
The middle setting has a formula. John Young's 1974 approximation gives the best time between checkpoints. It is the square root of two times the write cost times the mean time between failures. Daly's 2006 refinement adds terms that matter when a write is not small next to the interval. Converted to iterations for a Krylov run:
import math
def checkpoint_interval(write_s, mtbf_s, iters_per_s):
return int(math.sqrt(2 * write_s * mtbf_s) * iters_per_s)
rate = 2_564_096 / (48 * 3600) # 14.84 iterations per second
checkpoint_interval(30, 5_080, rate) # 8192
checkpoint_interval(30, 6 * 3600, rate) # 16892
Formula from Young (1974). The 30-second write cost is our assumption. The writeup times neither writes nor preemptions.
Run the formula backward to see what the choice says about the cluster. 8,192 iterations at 14.84 per second is 552 seconds. With 30-second writes, that interval is optimal when a preemption lands about every 85 minutes. That fits the post's "fatally preempted frequently." A pool that loses a node every six hours would want roughly 16,900 iterations instead. Measure your write time and preemption rate in the first hour, then set the interval from those two numbers.
The square root that overflowed
The last phase caused its own failure. The rational product reached about 1.76e11 bits, and "sqrt overflowed the mpz_t limb counter and aborted," the post reports. GMP stores the size of a big integer in a fixed-width field, and a number this large exceeded it. Devin rebuilt sqrt three times. The final version used GPU-accelerated NTT multiplication and finished in 88 minutes. The characters phase around it took 12 hours.
As a correctness check, the team factored a 344-digit number, C344, end to end. They did this while the RSA-260 Krylov phase was still in progress. It came out as a 136-digit prime times a 209-digit prime. A pipeline that factors a number you can verify independently is a far better test than any unit test on one stage. The check also cost almost nothing next to the main run.
The steering log
Lu reports that he "drove several parallel Devins," with an average of three and a maximum of 18 concurrent sessions over the three weeks. His own side of the conversation was 82,702 words across 3,328 messages in 192 sessions. That is about 17 messages per session and 25 words per message. He steered with short, frequent corrections. Agents themselves started 101 child sessions, and 36 of those ran with no human message at all.
He describes his job as supplying "executive function." He kept the goal hierarchy straight. "Recognizing when Devin was doing something unproductive and redirecting ('you don't need to take that measurement')" was part of that job. He also spotted repeated inefficiencies ("you can amortize this setup work") and pointed at directions nobody had tried. He writes that progress came fast enough that "I (and even Devin itself) could barely keep up."
Set those numbers beside the phase table. Agent sessions wrote the kernels, ran the queue and recovered from preemptions. Lu's messages decided which work to skip.
A three-week run offers thousands of plausible measurements, refactors and side quests. An agent with a good verifier will pursue all of them, because each one looks productive locally. Pruning that list took human judgment, and the run had one person to supply it. I made a version of this argument about short tasks in Stop Agents Building the Wrong Thing. RSA-260 shows the same failure mode over three weeks. There, each unpruned branch costs GPU-days instead of minutes.
Two harness controls keep a fleet like this bounded. Devin's v3 session API takes a parent_session_id, which keeps each child session attached to the run that started it. It also takes a max_acu_limit, which stops any single session from burning an open-ended budget:
curl -X POST "https://api.devin.ai/v3/organizations/$DEVIN_ORG_ID/sessions" \
-H "Authorization: Bearer $DEVIN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "Profile glas on one B200 workunit and propose one change with its measured rate.",
"parent_session_id": "'"$PARENT_SESSION_ID"'",
"max_acu_limit": 20,
"tags": ["rsa-260", "sieve-tuning"]
}'
Fields from the Devin v3 create-session reference: prompt is the only required one. Values are ours. Not run here.
You can write part of the pruning down. A rules file in the repository that the sessions share turns the corrections Lu quotes into standing instructions:
## Scope rules for long runs
- The current phase goal is in PHASE.md. Work not tied to it needs a one-line reason before it starts.
- Take a new measurement only if it changes a decision listed in PHASE.md.
- Setup that repeats per workunit goes into the image, not into each session.
- Every 30 minutes, report: what finished, what the oracle said, what comes next.
Our pattern, written from the corrections quoted in the writeup ("you don't need to take that measurement", "you can amortize this setup work"). The file cuts repeated corrections. A person still has to read the reports.
What it says about RSA
The run tells us little that is new about RSA, and the writeup says so. By Cognition's standard GNFS scaling, RSA-1024 at 309 digits is "merely 78x more computation than RSA-260," or roughly $30M. A cost of $30M is within reach of a well-funded attacker. That exposure is also why NIST's SP 800-131A transition disallowed 1024-bit RSA for new signatures back in 2013. "The fact that RSA-1024 is insecure is not news," Lu writes. RSA-2048 "remains roughly a billion times harder than RSA-1024," so this result does not affect the keys that protect most TLS today.
The result changes who can afford the attack. Cognition claims costs about 10 times lower than the previous public state of the art. The run used rented GPUs, with no specialized hardware and no CPU allocation from a national lab. The person who drove it says he "did not learn as much about NFS or GPU programming" as a specialist would have. If you still use 1024-bit RSA anywhere, the cost to break it is now in commercial range.
Worked example: the same sieve on 300 GPUs
Say you rerun the sieve on a smaller pool. Assume 300 preemptible GPUs instead of the average of 483 in the run, and a node lost every six hours on average. Keep ten-minute workunits and the $3.50 per GPU-hour from the writeup. The measured sieving bill of 3,813 GPU-days does not change with pool size. So the wall clock becomes 3,813 / 300 = 12.7 days, and the cost is 3,813 × 24 × $3.50 = $320,292.
Preemptions are the new line item. 300 GPUs × (12.7 days × 24 hours / 6 hours) is about 15,250 preemptions over the run. Each one loses on average half a workunit, or five minutes. That is 1,271 GPU-hours, or 53 GPU-days, which adds 1.4% to the sieve. The cost stays small because the unit of loss is ten minutes.
The preemptions also set two parameters. Every preemption is a workunit that times out. So tasks.maxtimedout must sit well above 15,250, or the server aborts the factorization partway through. Set tasks.wutimeout to about twice the slowest workunit, 1,200 seconds here. Then the server does not resubmit a slow client that is still alive.
The linear algebra cannot shrink its unit of loss. That is why its checkpoint interval matters and the interval of the sieve does not.
When a long run breaks
| Failure | What you see | Fix |
|---|---|---|
| Preemption in the Krylov phase | The job dies and restarts from its last checkpoint | Set interval from Young's formula with measured write time and preemption rate |
| Silent corruption | A consistency check fires hours in, or a solution that does not divide N | Keep online checks and bwccheck on. Never pair rolling checkpoints with skip_online_checks=1 (bwc README) |
| Workunits timing out | "Exceeded maximum number of failed workunits, maxfailed=100" | Find the cause first. The README names adrange and qrange set too large. Then raise tasks.maxfailed |
| Clients vanish on a preemptible pool | The factorization aborts after 100 timed-out workunits | Raise tasks.maxtimedout above the expected preemption count before launch |
| The driver dies | Nothing is scheduled | Restart from workdir/NAME.parameters_snapshot.N (README) |
| A stage that only fails at full size | sqrt aborts on GMP's limb counter at about 1.76e11 bits (writeup) | Run a full-size factorization you can check, like the C344, before the real one reaches its last stage |
| Agents on side quests | GPU-days spent on measurements no decision needs | A scope rules file plus short human corrections (writeup) |
The order to build a long agent job in
- Build an oracle for every stage. Use a small end-to-end run that must produce the right answer, and one rate measurement on real work. Ship nothing without both.
- Make the unit of work small enough to lose. Time one unit. Size the range so that one unit runs about ten minutes. Set the resubmit timeout to twice the slowest honest unit.
- Make the driver restartable. Give it a named work directory and a snapshot to resume from. Test it once by killing the driver on purpose.
- Place checkpoints by measurement. Time a write. Count preemptions in the first hour. Compute the interval from those two numbers.
- Run verification beside production. Keep online checks on. Check every pair of checkpoints. While the real job runs, run a full-size problem with a known answer.
- Bound the fleet. Tie child sessions to a parent. Set a spend cap per session. Tag sessions so that you can sum their spend.
- Prune in writing and by hand. Put scope rules in the repository. Have one person read the reports and say what to stop.
Most coverage will lead with the 260 digits. The log is the part to study if you run agents. A three-week, 4,900 GPU-day computation ran with an average of three agent sessions and one person who typed short corrections. Every phase let a machine check the work. The same harness pattern applies to jobs that have nothing to do with primes.
I wrote about the general shape in Harness Engineering. RSA-260 is the most concrete public case of it working at scale that I have seen.
Keep reading