Repository navigation
perf(runner-shared): decode memtrack frames for module events in parallel - #578
Conversation
1663e23 to
b4b751a
Compare
Merging this PR will regress 2 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Memory | find_module_events[1000000] |
1.1 MB | 4.4 MB | -75.16% |
| ❌ | Memory | find_module_events_with_stacks[1000000] |
1.1 MB | 4.5 MB | -74.98% |
| ⚡ | WallTime | find_module_events[1000000] |
1,269.3 ms | 187.1 ms | ×6.8 |
| ⚡ | WallTime | find_module_events_with_stacks[1000000] |
1,305.6 ms | 193.8 ms | ×6.7 |
| 🆕 | WallTime | find_module_events_with_stacks[10000000] |
N/A | 1.1 s | N/A |
| 🆕 | WallTime | find_module_events_with_stacks[100000000] |
N/A | 10.9 s | N/A |
| 🆕 | WallTime | find_module_events[10000000] |
N/A | 1 s | N/A |
| 🆕 | WallTime | find_module_events[100000000] |
N/A | 10.5 s | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing cod-3819-parallel-module-event-scan (f18817f) with cod-3819-memtrack-investigate-8-minutes-spent-in-teardown (2f1f983)
Footnotes
-
6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
b4b751a to
caf6e0d
Compare
|
8a11ab9 to
e66e598
Compare
…llel Finding the Mapping, Fork and Exec events decoded the whole artifact through one streaming deserializer on one thread. The streaming reader copies every key, string and stack payload into a fresh allocation before serde sees it. Every frame the encoder writes is a self-contained zstd frame, so decode_module_events now splits the artifact into its frames and decodes them on the rayon pool. Each worker decompresses a frame into a reused buffer and decodes events from that slice, so values are read in place instead of copied out of a stream. Results are concatenated in artifact order. A truncated last frame cannot be split off and is still streamed, keeping the events before the cut. The runner maps the artifact instead of streaming it from the file. The Mapping/Fork/Exec filter moves to MemtrackEventKind::is_module_event, and a test checks the new decoder against a filtered full decode over every event kind, several frames and a truncated last frame.
…ule event bench With frames decoded in parallel, CI-sized artifacts are fast enough to benchmark again: add the 10M and 100M event sizes the streamed decoder was too slow for. The 100M artifact with stacks is about 5 GB on disk and is mapped, not held in memory.
…vents Decompressing a whole frame before filtering it ties a worker's memory to the frame's decoded size, which the event count per frame does not bound: a frame of captured stacks decodes to hundreds of MB, held by every rayon worker at once. Stream each frame through the existing event decoder instead, so a worker holds about one zstd window and one event. Frames are still decoded in parallel and events keep their artifact order. This trades speed for bounded memory: on a 32-thread machine the 100M event search takes about 1.85 s instead of 1.0 s, still well ahead of the single-threaded streamed decoder.
e66e598 to
f18817f
Compare
TLDR: Finding the module events decoded the whole memtrack artifact on one thread through a streaming deserializer. Every frame is a self-contained zstd frame, so frames are now decoded in parallel from an in-memory buffer (stacked on #576).
decode_module_eventssplits the artifact into its zstd frames and decodes them on the rayon pool. Each worker decompresses a frame into a reused buffer and deserializesMemtrackEventfrom that slice, so values are read in place instead of copied out of a stream first. Results keep artifact order.MemtrackEventKind::is_module_event.memtrack_readerbench, local walltime on a 32-thread machine:Review notes
decode_module_eventsnow takes&[u8]and returns aVec.