Repository navigation
Conversation
Merging this PR will degrade performance by 81.64%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | sleep 1 |
245.8 µs | 1,338.7 µs | -81.64% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing spike/cod-3440-memtrack-musl (3a766f1) with main (8c84fcc)
Footnotes
-
6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
ad8eca1 to
2d8d983
Compare
|
`self_exe()` resolves the binary that internal subcommands are re-invoked through, and the memory executor hands that same path to `sudo setcap <caps>+ep` so the capabilities land on the binary that is actually exec'd. Reading an environment variable there means anyone able to set one variable chooses which file receives CAP_SYS_ADMIN and CAP_BPF. The override exists for the tests, where `current_exe()` is the test harness and cannot dispatch a subcommand. Nothing in production sets it -- the doc comment justified it with a launcher scenario that has no caller. Putting it behind `cfg(test)`, constant included, removes the escalation path outright while keeping the tests working; a release build now always resolves `current_exe()`. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`codspeed memtrack` and `codspeed exec-harness` are a re-exec of this binary and share nothing with the runner, but they were dispatched at the bottom of `run()` -- after the profile config is loaded, after the API client is built, and after `DiscoveredProjectConfig::discover_and_load` walks the filesystem. That last one is the problem: the re-exec runs in the benchmark's working directory, which is the user's project. A malformed `codspeed.yaml` there aborts the subcommand, so a measurement fails for a reason that has nothing to do with the measurement, and a `--config` given to the outer run is not forwarded to the inner one to override it. Move them into `run_internal`, called right after `Cli::parse()`. The logger match loses its internal arms for the same reason it had them. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The released Linux artifacts are `aarch64-unknown-linux-musl` and `x86_64-unknown-linux-musl`, and nothing in CI built either: a break in the argp stub, in the kernel-header paths or in the aarch64 `-lgcc` link flag would have surfaced for the first time during a tag-triggered release. The throwaway spike workflow used to cover this and was deleted with the spike. Both legs build on a native runner, with no environment variables, which is also what keeps `.cargo/config.toml` honest -- it has to carry the whole recipe on its own. The assertions are `readelf`-based rather than a `file` string, since rustc emits a static-PIE for x86_64 musl and spells it differently from aarch64, and `codspeed exec-harness --version` / `codspeed memtrack --version` answer only if both CLIs really are linked in. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`self_exe()` resolves the binary that internal subcommands are re-invoked through, and the memory executor hands that same path to `sudo setcap <caps>+ep` so the capabilities land on the binary that is actually exec'd. Reading an environment variable there means anyone able to set one variable chooses which file receives CAP_SYS_ADMIN and CAP_BPF. The override exists for the tests, where `current_exe()` is the test harness and cannot dispatch a subcommand. Nothing in production sets it -- the doc comment justified it with a launcher scenario that has no caller. Putting it behind `cfg(test)`, constant included, removes the escalation path outright while keeping the tests working; a release build now always resolves `current_exe()`. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`codspeed memtrack` and `codspeed exec-harness` are a re-exec of this binary and share nothing with the runner, but they were dispatched at the bottom of `run()` -- after the profile config is loaded, after the API client is built, and after `DiscoveredProjectConfig::discover_and_load` walks the filesystem. That last one is the problem: the re-exec runs in the benchmark's working directory, which is the user's project. A malformed `codspeed.yaml` there aborts the subcommand, so a measurement fails for a reason that has nothing to do with the measurement, and a `--config` given to the outer run is not forwarded to the inner one to override it. Move them into `run_internal`, called right after `Cli::parse()`. The logger match loses its internal arms for the same reason it had them. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The released Linux artifacts are `aarch64-unknown-linux-musl` and `x86_64-unknown-linux-musl`, and nothing in CI built either: a break in the argp stub, in the kernel-header paths or in the aarch64 `-lgcc` link flag would have surfaced for the first time during a tag-triggered release. The throwaway spike workflow used to cover this and was deleted with the spike. Both legs build on a native runner, with no environment variables, which is also what keeps `.cargo/config.toml` honest -- it has to carry the whole recipe on its own. The assertions are `readelf`-based rather than a `file` string, since rustc emits a static-PIE for x86_64 musl and spells it differently from aarch64, and `codspeed exec-harness --version` / `codspeed memtrack --version` answer only if both CLIs really are linked in. Reported by Greptile on #531. Refs COD-3440 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3294b83 to
8e2d7e3
Compare
GuillaumeLagrange
left a comment
There was a problem hiding this comment.
@not-matthias can you do a first round of review 🙏 ? Let's ignore the LD_PRELOAD removal overhead for the exec-harness simulation, I'll spec this in a dedicated ticket and we'll do it before we merge and release this.
not-matthias
left a comment
There was a problem hiding this comment.
Overall pretty good, just a few comments on how to better structure the code/comments
And a few notes on how to best structure the PRs for reviewers:
- Try to keep the changes as minimal as possible
- Do not modify unrelated comments/code -> should be done in separate PRs for easier review
- Remove obvious/LLM-written/bloated comments (using the deslop skill)
- you should also review the PR before putting it into review
GuillaumeLagrange
left a comment
There was a problem hiding this comment.
See comments for details.
I also see that there are unresolved comments.
When you handle a review, all comments should either be
- resolved and handled properly (very important to not silently resolve comments)
- answered with either counter arguments or questions so the reviewer can either accept your arguments or open a discussion
Finally, there are 35 commits in the PR, you may want to re-do commits before the next round of reviews
GuillaumeLagrange
left a comment
There was a problem hiding this comment.
See comments for details.
I also see that there are unresolved comments.
When you handle a review, all comments should either be
- resolved and handled properly (very important to not silently resolve comments)
- answered with either counter arguments or questions so the reviewer can either accept your arguments or open a discussion
Finally, there are 35 commits in the PR, you may want to re-do commits before the next round of reviews
|
Yeah I didn't see them since I needed to click on the see more button, I will check If they are already resolved, and try to reduce the amount of commits by squashing where it's coherent |
96d9261 to
ac9a23c
Compare
| #[command( | ||
| version, | ||
| about = "CodSpeed exec harness - wraps commands with performance instrumentation" | ||
| )] |
There was a problem hiding this comment.
no need for this IMO, the command is hidden, and its version will simply follow the codspeed version, which we should actually update in tehc argo.toml. Same for memtrack actually.
There was a problem hiding this comment.
I removed it, but I'm not sure I understand what I should update, Pass the crates to the codspeed version?
8a97c67 to
2b4a5c0
Compare
73f7a63 to
b933455
Compare
|
Want your agent to iterate on Greptile's feedback? Start a greploop in Claude Code and it will work through the open comments and keep going until this PR reviews clean. |
exec-harness injected a shared library into every benchmark to drive valgrind's instrumentation from inside the child. That only works on a dynamically linked executable, so statically linked benchmarks were silently unmeasurable, and it forced the harness to ship a `.so` next to its binary. The instrumentation is now toggled in exec-harness's own process, around the spawn: valgrind propagates the state across `fork`/`exec`, so the child is measured without anything being injected into it. The preload library, its compatibility check and the build script that produced it all go away, and the integration constants become plain consts. Since the child no longer switches instrumentation on for itself, exec-harness runs pass `--instr-atstart=inherit` to valgrind; with `no`, the child starts uninstrumented and the measurement comes back empty. Entrypoint runs keep the previous default. BREAKING CHANGE: exec-harness no longer ships a preload library. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both binaries kept their argument parsing and dispatch in `main.rs`, where nothing else can reach it. Move each into a `cli` module of its own crate and leave `main.rs` as a wrapper that installs a logger and calls `run_cli`. Nothing changes for the standalone binaries, but the runner can now link either CLI and dispatch it in-process. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The runner downloaded `exec-harness` and `memtrack` from GitHub releases at the start of a run, pinned by version, and `setup --mode memory` installed memtrack with `cargo install`. That is a network round-trip on every run, a version matrix to keep in sync, and two more artifacts to release. Link both crates instead and expose them as hidden `codspeed exec-harness` and `codspeed memtrack` subcommands, re-executing the current binary where the runner used to invoke the downloaded tool. They are dispatched before any runner setup: the re-exec happens in the benchmark's working directory, where an unrelated `codspeed.yaml` would otherwise abort the measurement. The binary installer, the memory setup no-op and the memtrack tool status go away. BREAKING CHANGE: `exec-harness` and `memtrack` are no longer downloaded or installed separately; the runner binary carries them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`exec-harness` and `memtrack` were released as their own artifacts. Nothing downloads them any more, so they stop being release units: their apt build dependencies move to the runner, which now builds the vendored libbpf and elfutils, and memtrack is depended on with its default features so the bundled subcommand carries the tracker. The released Linux artifacts are musl, and memtrack could not be built for them: `libbpf-sys` vendors elfutils, whose `configure` looks for `argp`, `obstack` and `fts`, none of which musl ships, and Debian's `musl-gcc` runs with `-nostdinc`, so the kernel UAPI headers libbpf needs are out of reach. The recipe lives in the cargo config so a plain `cargo build --target <arch>-unknown-linux-musl` works: seed the autoconf cache for the three checks, add a declarations-only `argp.h` stub on `CPATH`, add the UAPI header paths back through the per-target `CFLAGS`, and link `-lgcc` on aarch64 for libbpf's outline-atomic helpers. `close_range` goes through the raw syscall, since `libc` only declares it for glibc. CI builds both musl targets on native runners and asserts the artifact is static with `readelf`. The bundled subcommands no longer take `--version`, and the memtrack benchmarks call memtrack through the runner instead of installing it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
memtrack runs as a subcommand of this binary, so the allocator set in its own `main.rs` no longer applies to it. Set mimalloc on the runner binary instead, which covers memtrack and the rest of the runner alike. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With memtrack bundled, `setcap` targeted the runner binary itself, so every `codspeed` invocation ran in glibc's secure-execution mode: `LD_*`, `TMPDIR` and similar variables were stripped before the runner could forward them to the benchmark, in every mode. Grant the capabilities to a copy of the binary in `~/.cache`, keyed by its ELF build id, and run `codspeed memtrack` from it. The copy is installed in one sudo call, since `~/.cache` can be root-owned, and copies of other builds are pruned once they are a day old. The runner keeps the user's environment and no longer holds the eBPF capabilities itself. The memtrack benchmarks call `codspeed memtrack` from PATH, so CI grants that binary its capabilities directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
b933455 to
3a766f1
Compare
exec-harnessandmemtrackare compiled intocodspeedand reached as hidden subcommands, andexec-harnessno longer injectslibcodspeed_preload.so. The two are one change: a staticallylinked musl binary cannot be preloaded into a glibc process, so the single binary was blocked on
removing the preload.
One release artifact instead of three, 11.0 MB compressed against 13.5 MB today. The download
machinery and both installer pins are gone.
Three things worth a reviewer's attention:
fork, so it also covers fork/exec/wait and the child's pre-
mainstartup. It is a constant~940k Ir per exec-harness benchmark, not a percentage — stable to 0.22% across a 1000× range of
benchmark size. Shipping as is: no forced baseline, no history surgery.
benchmark URI, since
LD_PRELOADis what used to be inherited by every descendant.setcapnow lands oncodspeeditself, so memtrack's five capabilities sit on the CLI.They are
+epwith no inheritable set, so a spawned benchmark does not receive them.Also removes the user-facing "CPU Simulation mode does not support statically linked binaries"
error: nothing is injected into the benchmarked executable any more.
Closes COD-3218
Closes COD-3440