[ ULTRO · LAB ]

The Debris Field

A manifesto on clean data.

The debris field, four passes onRest
Reading 02 / The same corpus, sampled at every remove

In 1978, a NASA scientist named Donald Kessler ran the numbers on orbit and did not like the answer. Every satellite launched adds a small amount of debris: paint flecks, spent stages, the odd bolt. Debris travels at orbital velocity, so even a fleck can crack a solar panel or a window. Above a certain density, Kessler showed, collisions start making more debris than launches do. The belt feeds itself. Past that point, adding fewer satellites does not fix it. The junk is now the primary source of junk.[1]

Nobody clears an orbit. There is no broom for low Earth orbit and no plausible one. Once the cascade starts, the only real choice left is where you still fly clean.

The internet is running the same experiment with words instead of aluminum.

For twenty years, every large language model trained on the same rough material: the crawled, scraped, sold, and stolen text of the open web. That web was mostly written by people, which is why the first generation of models sounds like a person. It will not stay that way. In 2024, a team of researchers gave the failure mode a name: model collapse. Train a model on its own output, then train the next model on that model's output, and repeat. Measured quality falls with each pass. Rare facts vanish first. Variance narrows. The model converges on an average of an average of an average, and the average gets duller every round.[2]

That paper studied a laboratory version of the problem. The open web is running the live version, at scale, right now. Generated text is cheap to produce and is being published faster than any population of humans could type. A growing share of what a 2027 crawler indexes will be the output of a 2025 model, edited by a 2026 model, summarized by a 2024 model. The next generation of models will train on that. Debris breeding debris, in orbit and on disk.

A reasonable instinct says: hoard what is still clean. Archive everything written before the flood, and train on that instead.

This is where low-background steel comes in. Steel made after 1945 carries faint radioactive contamination from atmospheric nuclear tests, absorbed straight out of the air during smelting. Instruments sensitive enough to detect a single particle of dark matter cannot tolerate that contamination, so physicists go looking for steel forged before the first bomb. Most of the world's supply now comes from ships that sank before 1945 and have sat on the seafloor ever since, cut and sold a hull at a time.[3]

Low-background steel is a fixed, shrinking pool. Nobody is making more of it. Every year a little more of it gets used up in detectors, and the remaining ships get harder to reach and more expensive to salvage. It is clean, and it is finite, and it describes 1944, not now.

Pre-2022 internet text has the same shape. It is comparatively clean, and it is a fixed, aging archive of a world that keeps moving. It has no record of a browser agent booking a flight, no record of a warehouse worker overriding a scanner, no record of an analyst catching a bad number before it reached a client. Those processes did not exist yet, or nobody happened to write them down while they were happening. Hoarding old text is not a growth strategy. It is a museum.

Here is the part that should worry any team building agents: the internet was never a record of work. It is a record of conclusions. People write down the decision, the summary, the finished report. Almost nobody writes down the fifteen minutes of judgment calls that produced it. The hesitation before approving the invoice. The second look at the chart before the trade. The sequence of clicks an experienced operator uses without thinking, that a new hire takes six months to learn by watching over someone's shoulder.

That is exactly the material an agent needs to be trained on, and it does not exist in any archive, clean or otherwise, because nobody was recording it. It has to be made.

So we make it. We design a real experiment around a real process, put trained operators in front of it, and instrument the whole thing: video on their hands and their face, audio on the room, a capture of the exact screen, a log of every keystroke, sensor data where it exists, the documents they touch, a transcript of anything said out loud. One clock ties every stream together to the millisecond. Every person in the room signed a form before any of it started. Every file has a custody record from the moment it was captured to the moment it left our hands. That is production discipline, not scraping, and it is the only way this material gets to exist at all.

Provenance is the product.

A dataset without a clean chain of custody is a rumor. A dataset with one, naming who did what, when, on what device, with what consent on file, is an asset a company can put a number on and defend in a room full of lawyers. That is the difference between data you found and data you can stand behind.

Ownership is where the argument closes. A model you rent expires at the end of the license term and every competitor can rent the same one. A dataset built from your own operation does not expire, cannot be rented by anyone else, and gets more valuable every year the open supply gets worse. Companies that treat their own ground truth as an asset, not an afterthought, will compound an advantage that a better prompt or a bigger model cannot close. Everyone else will keep drawing from the same shrinking, increasingly contaminated well.

None of this argues against synthetic data. Synthetic generation is fast and cheap and useful for exactly what it is: an amplifier. It is only ever as good as the real examples it was seeded from, and it inherits every gap and every bias in that seed without telling you. Somebody still has to make the seed, from real people doing real work, before any of the multiplication starts.

That is the work. Design the experiment. Put real people in it. Record every modality on one clock. Document the chain of custody without a gap. Hand over the only copy. Do it again for the next process, and the one after that, while the open web keeps filling with its own exhaust and calling it content.

Backed by