# The flag in an hour, after eight and a half hours of nothing

> The L3AK "Breadcrumbs" challenge, an AD1 image, a 7 GB memory image and a packet capture, worked by the swarm twice. The first run could not read the memory image at all and gave up its first segment in twenty-three minutes; it factored the TLS server certificate two hours and sixteen minutes in and still stopped without the flag after eight and a half hours. The second run, after I changed the platform and nothing else, read the memory image from the first minute, factored the same certificate in nine, read a protected key out of process memory, and finished with the complete flag in about an hour.

By Halil Öztürkci, 6 October 2026. https://dfirswarm.ai/blog/breadcrumbs-the-flag-in-an-hour-after-eight-hours-of-nothing

"Breadcrumbs" is an L3AK challenge from the community compilation at [Azr43lKn1ght/DFIR-LABS](https://github.com/Azr43lKn1ght/DFIR-LABS) (MIT; the repository asks that its author and the labs be mentioned, and this is that). It is rated Insane, and it asks no questions. It gives three files of three different shapes, says that a workstation started acting strangely and slowed down, and asks for one thing: the flag, in the format `L3AK{...}`, somewhere in the evidence. The three files are a 346 KB AccessData AD1 logical image, a 7.0 GB Windows memory image, and a 14 MB `pcapng` capture.

I ran it twice. This post is about the two runs together, because the difference between them is the whole point. The first run (1 October) could not read the memory image at all, wrote its own reader for the AD1, ran into a closed door on the one route that mattered, and stopped after eight and a half hours with a candidate flag it could not stand behind. Between the two runs I changed the platform, not the prompt: I baked the exact Windows kernel symbol table into the memory image, added an AD1 recipe, put a few missing programs in the job images, and landed some run-safety fixes. The second run (6 October) finished in about an hour with the complete flag, no hint from me, and not one request to the operator.

As with the earlier posts, this one names the methods and none of the answers. The flag is sealed. Where a sentence needs a value to make sense I use a placeholder: "the request-path prefix" for the part of the flag that was already a URL path in memory, "the short clue" for the small text object the server returned, "the decrypted suffix" for the last component, "the server key" for the TLS private key the swarm recovered, and "the protected key" for the credential-protected material it read out of process memory. Credential-handling steps I describe only at the level of what an examiner looked at, in the defender's voice; this is a CTF image and there is no attack recipe here. Where an agent's chosen name would give part of an answer away I would mask it, but here the names are about the evidence shapes, not the answer, so I use them. The figures are in the box beside this text.

One caveat up front, because it governs everything below. This is one run per round, not a controlled experiment. Between the two runs the images, the symbols, the AD1 support, the packs and the harness all changed together, and I knew the first run when I set up the second. The challenge has no answer key that I can check against: the upstream repository ships no solution for Breadcrumbs, and I found no write-up. So when I say the swarm got the flag, I mean it established it from the evidence, with two independent reconstructions that agree byte for byte; I do not mean it was graded correct by the challenge platform, because there is nothing to grade it against. And the examination as a whole is limited: the flag is settled, but how the intrusion started, whether anything persisted, and what actually caused the slowdown stayed at the edge of what three snapshots can show.

## The setup

The conditions of the run, once. Every agent reasoned in a microVM of its own with no forensic programs in it, and every extraction and parse ran as a job in a throwaway worker VM, its output sealed into the store and cited by job id. The team mixed `gpt-daybreak-blue`, `gpt-6.1-sol` and `gpt-6-luna` through a subscription, and nobody was assigned anything. There was no token cap and no wall clock; spend was recorded and shown, and nothing stopped for it. The run was set to go on until every question in scope had an answer that stands, with a regroup after fifteen quiet minutes. One thing was new to this case and matters later: the goal marked the flag question as one the run must establish. Under that rule the run cannot end on the flag with a reviewed limitation the way an ordinary question can; a limitation is not an acceptable disposition for it. That rule exists because of what happened the first time.

![The console's story tab early in the second run: the kickoff's two catalogue announcements, the first agents introducing themselves, and the first claims being negotiated](https://dfirswarm.ai/assets/img/blog-breadcrumbs-opening.webp)
*The opening of the second run, read back from the board: the harness has already catalogued the AD1 and the capture, and the seats are splitting the three evidence files between them.*

[Video (22.0 MB, MP4): A silent 75-second replay of the second run's board, scrolled from its first minute to the finish, with a clock in the corner; the flag's fragments are masked](https://dfirswarm.ai/assets/video/blog-breadcrumbs-replay.mp4?v=e1cd4ddfef)
*The whole second run's board, scrolled from minute 0 to the finish in 75 seconds. The clock in the corner is the minute of the run; the flag's fragments are masked.*

![The console's agents tab at the end of the second run: one microVM per agent on a single image digest, each put away with its disk kept, and the finish line certified eight of eight](https://dfirswarm.ai/assets/img/blog-breadcrumbs-vms.webp)
*The agents tab at the stop: one microVM per agent, all on the same base image digest, each put away with its disk kept, and the finish line certified eight of eight.*

## The first run: a memory image the platform could not read

The first run started the way the second would, with the team splitting the three files, and then it hit two walls at once.

The memory image was the larger wall. Volatility needs a kernel symbol table that matches the exact Windows build in the image, and it fetches that table from Microsoft's symbol server over the network. These VMs have no network, by design. The kickoff's own catalogue pass tried to read the memory image and failed on precisely this: it detected a Windows image, found the kernel, computed the symbol identifier it needed, and then could not reach `msdl.microsoft.com` to download the matching table, so it produced nothing. Every memory plugin an agent tried for the next hours failed the same way. The network side of the case could see a 14 MB transfer and eleven TLS sessions but could not read a single process out of the 7 GB image.

The AD1 was the smaller wall. The Sleuth Kit does not read AccessData's AD1 format, and the platform had no recipe that did, so the catalogue marked `disk.ad1` "not catalogued." A seat wrote its own reader: it decoded the AD1's compressed streams by hand, and a critic reproduced one of the recovered files at the same offset with the same hash. That worked, but it was work the platform should have done for free.

Memory strings still held the shape of the flag. A literal scan found the request-path prefix repeated across the image as a URL to the server, and plaintext GET lines for it, next to URLs for the other transferred files. But the prefix in memory stops where it stops. It has no closing brace. One seat wrapped it in the required `L3AK{...}` syntax and offered that as the flag; the critic would not let it stand:

> **Objection · Source Critic (s3472f005), gpt-6.1-sol · 9 min in · #57**
>
> Source-first Q1 review: disputed E-47's exactness inference, not literal memory body. … CASE.md wrapper specifies syntax, not correctness or body completeness. Repeated URL copies may share one incomplete/lure source. Q1 must expose this unestablished part unless another artifact discriminates; do not silently claim byte-reproduction of an added brace.

That was the right call, and it is the whole first run in miniature: the swarm could see where the flag began and could not reach the rest, and it refused to pretend otherwise. Twenty-three minutes in, with every question answered only as far as the evidence reached, the report author published, the reviewers signed, and the run ended itself. Under its stop policy a reviewed limitation is a legitimate way to finish, so the first segment closed `examination_limited`, with no flag. That exposed a gap in the harness: there was no way for me, as the operator, to say "a limitation is not acceptable for this one question" to a flag-only challenge. I resumed the run with exactly that demand, by hand, and it fought the tooling the whole way (the operator role was not enrolled, so my continuation question did not reach the register until I added it a second way).

What the resumed run then did is the reason the second run was worth setting up. It could not read the memory image, so it went at the network capture's cryptography instead. The eleven TLS sessions all used a non-forward-secret RSA cipher suite, which means a server with a weak key is decryptable after the fact. A seat parsed the server's certificate, found the RSA modulus suspiciously structured, ruled out the easy mistakes, and then factored it:

> **Breakthrough · PacketCartographer (s3472f001), gpt-daybreak-blue · 2 h 16 min in · #382**
>
> L43 will explicitly test the source-quality rival. The captured DER parses consistently in two independent loaders and, more decisively, its exact Fermat factors derive a private key that self-checks against that DER and causes tshark to authenticate/decrypt coherent HTTP requests/responses across the retained capture.

Here is the detail I keep coming back to. The harness proposes a stop when nothing has moved for a while; it had proposed one, because the factoring had not landed yet. The stop proposal was written at 10:01:32.446Z. The factoring finding was recorded at 10:01:33.924Z. One and a half seconds apart. I had the choice to end the run one breath before its single biggest result, and the only reason I did not is that my standing instruction on this run was to let it keep going. The decrypted capture gave up five HTTP objects: a directory listing, a missing favicon, the short clue, a small protected blob, and a 14 MB unsigned executable. It was real progress, two and a quarter hours in, on the back of a cryptographic attack the team ran itself. And it still was not the flag, because the last component was locked inside that executable and the key to it was in the memory image the platform could not read. I stopped the run at eight and a half hours.

The first run also had two incidents I told straight in my notes and will tell straight here. The custody chain flagged a change to the evidence folder; the cause was mine, not the swarm's. A download helper I was running wrote a checksum file into the live run's evidence directory, and the agents' integrity sweep caught it within two seconds and stopped trusting the inputs until I adjudicated. It was an added file, not a change to any evidence object, and when I re-hashed all four files they matched the kickoff manifest exactly. The lesson was mine: never point a helper at a live run's source folder. Second, at the stop two trace lines had been written while the collector was briefly down and ended up outside the hash chain, spilled to a side file; the stop chained them, but a clean run should have none. Both went on the fix list.

## What I changed, and nothing else

Between the runs I did four things, all to the platform.

I baked the Windows kernel symbol table into the memory image. I pinned the exact symbol file the image needs by its public identifier and hash, converted it to Volatility's format at image-build time, verified the result against a canonical hash, and placed it where Volatility looks. The build never fetches it during a run; the operator accepts Microsoft's symbol terms once, at the host, and the table travels inside the image. The honest sentence I keep on this: Microsoft's terms grant use for debugging and testing your own software, and whether examining a third party's memory image falls inside that is not a question I have settled.

I added an AD1 recipe to the base pack, so the catalogue reads AccessData logical images the way it reads disk images and archives, instead of leaving a seat to decode the container by hand.

I put the missing programs into the job images, so the second run would not need me to hand it tools mid-flight the way the first did (the first run needed me to supply `steghide` and to build `aeskeyfind` for the image's architecture and hand both in as operator material; neither was in any image).

And I landed the run-safety fixes the first run had asked for: a read-only, non-executable copy of the evidence so a host helper can never touch the live inputs; an advisory token alert; an enrolled operator identity so a continuation question is admitted; and the ability to mark a question as one the run must establish rather than bound with a limitation. That last one is the direct descendant of the first run's twenty-three-minute finish.

I changed no prompt, gave no hint, and did not intervene once the second run started. The only operator action in the second run's whole trace is the harness's own `stop` after the finish.

## The second run, minute by minute

The second run started at 05:48 local. By the end of its first minute the split had formed the way it always does, from pivots rather than a plan. FlagPath took the flag itself and said it would start from the AD1 and follow the artifacts into the capture or memory. Packet Cartographer took the network. Cross-Source Chronologist took the timeline, and when it found another seat had already claimed it, Timeline Weaver pivoted to the workstation's processes. Workstation Analyst took the disk and static-analysis side. Incident Chronologist, finding the timeline taken, became the report author. The three `gpt-6-luna` seats became the independent reviewers: Memory analyst, Memory artifact analyst, and Capture Analyst.

The contrast with the first run showed in the first ninety seconds. The harness's catalogue had already read both the memory image and the capture before any agent did anything, and the AD1 was read completely within the first minute: the recipe pulled all 122 items out of the logical image (72 files and 50 folders, every file hash matching the container's own metadata), where the first run had a seat decoding streams by hand. And Timeline Weaver's very first memory jobs just worked. Here is the sequence from the trace, by the clock:

- **02:49:52**: the first `vol --offline` call, reading plugin help, on the baked symbol table.
- **02:50:12**: `windows.cmdline` and `windows.pstree` running against the 7 GB image, offline, no network.
- **02:50:38**: the process tree is already a finding, `7za.exe` from a `Downloads\x64` folder spawning a browser and a console host, with a burst of some twenty-odd orphaned browser processes around it.

The thing that had closed the first run's door for hours was now a thirty-second formality. Volatility worked from the first minute, and the memory side of the case was open.

![The console's ledger tab in the first minutes of the second run: the AD1 extraction, the memory image's process tree and the capture's summary already recorded as findings](https://dfirswarm.ai/assets/img/blog-breadcrumbs-firstminute.webp)
*The ledger in the first two minutes: the AD1 read completely (E-1), Volatility parsing the memory image offline (E-2, E-8) and the capture summarised (E-4). In the first run no memory finding could exist.*

![A two-lane timeline diagram: the first run's lane with "memory unreadable" spanning hours, "AD1 read by hand," "first segment stops, no flag" at 23 minutes, and "certificate factored" at 2 h 16 min; the second run's lane with "Volatility working" at minute 2, "certificate factored" at minute 9, and "complete flag" at 58 minutes](https://dfirswarm.ai/assets/img/blog-breadcrumbs-timeline.webp)
*The two runs on one clock. The marks are from the runs' own timestamps. In the first run the memory image was unreadable throughout and the certificate fell at two hours sixteen minutes; in the second the memory image was readable from minute two and the same certificate fell at minute nine.*

### The false key in memory, run down and rejected

Before the good route, the honest dead end. The capture's TLS used a non-forward-secret RSA suite, so the team knew from the second minute that a server private key would unlock everything. Two of the reviewers went looking for one in memory. A full-image keyword scan reported a hit on the text `BEGIN RSA PRIVATE KEY`, and for a few minutes it looked like the key was simply lying in the dump. It was not. The first seat to raise it was also the first to take it back, on a bounded re-scan of the surrounding bytes:

> **Objection · Memory artifact analyst (sd2925207), gpt-6-luna · 4 min in · #54**
>
> Correction to my prior lead: E-51 supersedes E-37. Independent bounded scan j000027 confirms only undelimited text `BEGIN RSA PRIVATE KEY` at offset 2732212459 plus base64-like context; it is NOT a complete PEM marker. I retract suggestion it may be directly parseable; key hypothesis remains unverified.

The report author saw a second seat still treating the same offset as a candidate and stepped in to hold the line before anyone built on it:

> **Kept out · Incident Chronologist (sd2925205), gpt-6.1-sol · 5 min in · #57**
>
> Before assuming candidate is PEM, please see critic #54/E-51: independently scanned bytes found undelimited needle text, not dashed delimiter at offset 2732212459. Offline parsing/public-key comparison remains useful discriminator; no private-key inference until validated.

The second seat then corrected its own entry too, superseding its candidate finding with a plain statement that the marker was undelimited text and no usable key. The whole episode is on the record as three ledger entries that supersede one another, and nothing downstream ever cited the false lead. It is the same marker the first run had chased; the difference is that the second run disposed of it in nine minutes with the real route already running beside it.

![The console's story tab at five minutes: a seat retracting its own lead about a private-key marker in memory, the capture seat saying it will use a candidate only if it matches the certificate, and the report author telling the others not to build on it](https://dfirswarm.ai/assets/img/blog-breadcrumbs-falselead.webp)
*Five minutes in: the memory analyst corrects its own lead twice, the capture seat says it will use a candidate key only if it matches the certificate, and the report author tells the others to wait for validation.*

### The certificate, factored at minute nine

The real route was the one the first run had found at two hours sixteen minutes. Packet Cartographer gave up the memory-key idea, parsed the server's certificate out of the capture, and looked at the modulus itself. It was an odd, oddly sized RSA modulus with its two prime factors close together, which is the textbook condition for Fermat factorization. The seat factored it on the first iteration, derived the private key, checked the key against the certificate, and handed it to TShark to decrypt the whole capture:

> **Ledger E-62 (Packet Cartographer, sd2925201), 9 min**
>
> The TLS application data was successfully decrypted by deriving the server RSA private key from the certificate's trivially factorable modulus, validating the key, and supplying it to TShark. The workstation issued HTTP GET requests to the server for `/` (200, directory listing), `/favicon.ico` (404), the short clue (200), a small protected blob (200), and a 14 MB unsigned executable (200).

Every step of that is in a sealed job: the factorization proof in one job, the decryption and the five exported objects in another, each cited by id. The private key itself is withheld; the entry is marked sensitive and the key never reached the board, the ledger or a command line. The five objects matched, to the byte, the five the first run had recovered two hours in. My reading, which one run cannot prove: minute nine against two hours sixteen is not a smarter swarm but the same swarm with a memory image it could read, so it could spend the network seat's attention on the certificate instead of hunting for a key that was never going to be in the dump.

![A diagram of the evidence chain with the answers redacted: three evidence files feeding into (a) the AD1 recipe, (b) Volatility over the memory image, (c) the TLS certificate factored into a server key that decrypts the capture into five HTTP objects; the five objects feeding the injector DLL, the PyInstaller executable, and the protected blob; and a credential read from LSASS process memory feeding the final decryption; the three recovered components joining into one sealed flag](https://dfirswarm.ai/assets/img/blog-breadcrumbs-chain.webp)
*The chain of evidence as the second run built it, with every recovered value redacted. Three sources, one certificate break, one credential read from process memory, three components that join into the flag.*

### What the seats did with the memory image

While the network seat broke the capture, the two sol-model seats reconstructed the workstation from memory, and this is where the second run earned its "readable from minute one."

Timeline Weaver scanned the memory image's file objects and recovered two cached binaries from the same `Downloads\x64` folder: the `7za.exe` and, beside it, a custom `cryptbase.dll`. That DLL is the heart of the case. Static disassembly of its startup code showed it was not the system library it named but a custom injector: it creates a browser process suspended, decodes an embedded resource with a repeating-key XOR, and writes that into the suspended process with the usual allocate-write-create-thread sequence. The seat decoded the resource offline, in a job, by re-implementing the XOR rather than running anything, and out of it came shellcode and a second embedded DLL carrying the server's URLs. Workstation Analyst took that embedded DLL and read its behaviour statically: it downloads the protected blob and the short clue into a user AppData folder and saves the executable into the per-user Startup folder. It also noticed, and kept, a small discrepancy: the URL the DLL is configured to fetch is spelled with an extra letter compared with the filename seen on the wire. That is exactly the kind of detail a report should preserve rather than smooth over, and the seat flagged it rather than equating the two.

None of this was run. Everything came from reading bytes: the process tree from Volatility, the cached DLLs from file-object scanning, the injector's logic from a disassembler, the embedded payload from a decoder written in the job. An agent whose own VM holds no forensic tool reconstructed a DLL-sideloading injector and its downloader purely by static examination, and a second seat re-derived the decisive parts.

### Reading the executable without running it

The 14 MB executable the capture delivered was a PyInstaller bundle, and the last piece of the flag was inside it. FlagPath parsed the PyInstaller archive statically and pulled out its main Python module as an inert marshal stream. To read that stream without executing it, the seat forged a tool in the moment:

> **Handed on · FlagPath (sd2925200), gpt-daybreak-blue · 18 min in · #157**
>
> Forged tool `marshal_inspect` v1 (python3): Statically parses a Python marshal stream (including 3.11+ code objects) into JSON metadata and bytecode without executing it. … It is in your tool list after your next inbox or wait.

![The console's story tab at eighteen minutes: FlagPath announcing the forged tool marshal_inspect, which parses a Python marshal stream without running it](https://dfirswarm.ai/assets/img/blog-breadcrumbs-forge.webp)
*Eighteen minutes in: FlagPath announces the tool it forged to read the executable's Python module as data, while the other seats review their answers and the catalogue keeps cataloguing derived objects.*

With that tool it read the module's logic as data: a routine that reads the small protected blob, hands it to Windows' data-protection API, uses the returned value as an AES key, and decrypts an embedded ciphertext with an embedded IV. The seat was careful to record something subtle about the bytecode: the routine's own final instruction returns the wrong object, so the decrypted bytes are computed and then not returned. Reading that statically is worth more than running it would be, because it tells you the author's intended transform exactly, and it never puts adversary code in motion.

### The protected key in process memory

The last component needed the AES key, and the AES key was protected by Windows' per-user data-protection system, whose master key lives in process memory. This is the credential-bearing step, and I will describe it only as examination method. FlagPath dumped the relevant process's memory with Volatility, parsed the data-protection state out of that dump in a clean-room job, used it to unprotect the small captured blob, and got the key the executable's routine expected. That key decrypted the embedded ciphertext to the final component. Every intermediate value in that chain, the process memory, the unprotected material, the key, the plaintext, lives only in sealed, sensitive job outputs; a redacted hand-over package removes what they cite, and nothing of them appears on the board or in the report beyond the statement that the step happened.

### Three components, assembled and reproduced

The flag was three pieces from three places: the request-path prefix recovered from memory, the short clue recovered from the decrypted capture, and the suffix decrypted from the executable with the protected key. FlagPath assembled them in a declared-scope job and recorded the result. Then the run did the thing I most want it to do: two other seats reproduced it independently, from the sealed sources, without trusting FlagPath's assembly.

> **Handed on · Packet Cartographer (sd2925201), gpt-daybreak-blue · 44 min in · #195**
>
> L-17 independent Q-1 verification completed. E-124 records a trusted static reconstruction over sealed refs: … yields the exact flag. Recovered code was never executed; intermediate DPAPI/AES key remains withheld.

Capture Analyst did it a third time, from the network side, recomputing the AES step from the sealed key material and reassembling the components. Three reconstructions, the same bytes, the same digest, two of them by seats that did not do the original recovery. The report author folded it into a standing answer, the two reviewers recorded their attestations as "established," and nothing disputed it.

## The operator who never spoke

Here is the plainest measure of what changed. In the first run I was in the loop constantly: I answered a clarification, I argued about whether a formatted candidate counted as an answer, I supplied `steghide`, I built and supplied `aeskeyfind`, I adjudicated a custody alert I had caused myself, and in the end I stopped the run by hand. The operator-request log for that run is long.

In the second run the operator-request count is zero. No clarification, no tool handed in, no network opened, no hint, no intervention. The swarm had everything it needed inside the images, and the only thing that bears my name on the whole trace is the automatic `stop` the harness fired after the finish sealed itself. The feature that would have let me force the first run past its twenty-three-minute limitation, the must-establish rule, was in the goal the second time and the run never needed me to invoke it, because it simply established the flag.

![The custody tab at the end of the second run: a clean verdict, the finish line certified eight of eight, and the run marked done with one agent unfinished](https://dfirswarm.ai/assets/img/blog-breadcrumbs-custody.webp)
*The second run at the stop: evidence re-hashed and unchanged, every chain intact, the finish line certified, and the whole thing sealed without an operator in the loop. One seat was still mid-write when the finish landed.*

## The finish

The ending had one wrinkle worth telling, because it is the harness, not the swarm. When the report author first called `done`, the goal's checks refused it: the flag answer leaned on a few entries that recorded failed exploratory jobs, and the gate would not accept failed routes listed as if they bounded the successful one. The fix was bookkeeping, not evidence: the author re-recorded the answer without the irrelevant failed-route entries, the reviewers refreshed their attestations onto the new version, the summary and narrative were repointed, and the next `done` went through. The finish line then certified structure, citations, attestations and unchanged inputs, eight of eight. It does not certify correctness, which is why "established" and "verified" are different words throughout this post. At the stop the host re-hashed the four evidence files, found them byte-for-byte unchanged, verified every chain, and sealed a draft release. The whole second run was about an hour of work and a few minutes of sealing.

## What it found and what it missed

Graded against the evidence, by the question the goal actually asked:

| Question | What it asks | Result | Why, in method terms |
| --- | --- | --- | --- |
| Q1 | The flag, exactly, with every step to reproduce it | Established from the evidence | Three components recovered from three sources, assembled, and independently reproduced twice from sealed outputs; not verified against the challenge platform, because none exists. |
| Q2 | What happened on the workstation | Partial | The archiver, the injector DLL and the downloader are recovered and read statically; initial access, successful persistence and the measured cause of the slowdown are at the limit of three snapshots. |
| Q3 | The network: hosts, protocols, transfers, ties | Partial | The capture is fully decrypted and every transfer identified and tied to memory; which process owned the socket, and whether this was interactive command-and-control, the evidence cannot settle. |
| Q4 | Cross-source timeline and lessons | Established within the acquisitions | A merged timeline that keeps each source's clock separate, with the inconsistency between the kernel clock and the process times preserved rather than smoothed. |

The finish line certifies that the record holds together. It does not say the answer is the true one, and on this case there is no ground truth to appeal to. What I can say is that the flag rests on two independent reconstructions that agree to the byte, and that the parts I have called partial are called partial because the evidence stops, not because the swarm did.

## What I take from this run

Running the second time took about an hour, 148.7 million tokens and roughly 1,500 model calls, on a subscription with no metered bill. The first run took eight and a half hours and nearly 790 million tokens and never reached the flag. That is roughly an eighth of the time and a fifth of the tokens, with the flag, but I will not dress it up as a benchmark: it is one run against one run, and the thing that changed is not the swarm, it is what the swarm could read.

That is the lesson I actually trust from putting these two runs side by side. The intelligence was there in the first run. It wrote its own AD1 reader. It factored an RSA certificate out of a packet capture and decrypted the whole thing. It refused to call a prefix the flag when the brace was not in the evidence. What it lacked was a platform that let it read a 7 GB memory image without the internet, and that one missing table turned a one-hour case into an eight-and-a-half-hour failure. Most of what makes a swarm like this useful is not the model. It is whether the evidence is readable at all, offline, inside the box, from the first minute. Bake the symbol table in, teach the catalogue the file format, put the tools in the image, and the same agents walk through a door that was closed to them a week earlier.

I will also say the unglamorous part. The night before this run, a smaller smoke run on a different, easier case came out weaker than its own baselines, for reasons I have not fully pinned down. So I am not claiming the platform got uniformly better; I am claiming that on this case, this time, the change I made was the change that mattered, and the record shows why. The record of both runs, every board post, every ledger entry and every sealed job, including the first run's closed door, is what this post is written from.
