Skip to content

feat!: ship one binary — bundle exec-harness and memtrack, drop the LD_PRELOAD hack - #531

Open
moha-bekh wants to merge 6 commits into
mainfrom
spike/cod-3440-memtrack-musl
Open

moha-bekh wants to merge 6 commits into
mainfrom
spike/cod-3440-memtrack-musl

Conversation

@moha-bekh

@moha-bekh moha-bekh commented Sep 7, 2026 •

Copy link
Copy Markdown
Member

exec-harness and memtrack are compiled into codspeed and reached as hidden subcommands, and
exec-harness no longer injects libcodspeed_preload.so. The two are one change: a statically
linked musl binary cannot be preloaded into a glibc process, so the single binary was blocked on
removing the preload.

One release artifact instead of three, 11.0 MB compressed against 13.5 MB today. The download
machinery and both installer pins are gone.

Three things worth a reviewer's attention:

  • The measurement baseline shifts. The measured region now starts in exec-harness before the
    fork, so it also covers fork/exec/wait and the child's pre-main startup. It is a constant
    ~940k Ir per exec-harness benchmark, not a percentage — stable to 0.22% across a 1000× range of
    benchmark size. Shipping as is: no forced baseline, no history surgery.
  • Attribution rests entirely on the spawn-chain walk. Children no longer self-label with the
    benchmark URI, since LD_PRELOAD is what used to be inherited by every descendant.
  • setcap now lands on codspeed itself, so memtrack's five capabilities sit on the CLI.
    They are +ep with no inheritable set, so a spawned benchmark does not receive them.

Also removes the user-facing "CPU Simulation mode does not support statically linked binaries"
error: nothing is injected into the benchmarked executable any more.

Closes COD-3218
Closes COD-3440

@codspeed

codspeed Bot commented Sep 7, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 81.64%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

❌ 1 regressed benchmark
✅ 30 untouched benchmarks
⏩ 6 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation sleep 1 245.8 µs 1,338.7 µs -81.64%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing spike/cod-3440-memtrack-musl (3a766f1) with main (8c84fcc)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@moha-bekh moha-bekh changed the title COD-3218 / COD-3440: remove the LD_PRELOAD hack and port exec-harness to musl feat!: ship one binary — bundle exec-harness and memtrack, drop the LD_PRELOAD hack Sep 17, 2026
@moha-bekh
moha-bekh force-pushed the spike/cod-3440-memtrack-musl branch from ad8eca1 to 2d8d983 Compare September 17, 2026 14:33
@moha-bekh
moha-bekh marked this pull request as ready for review September 17, 2026 14:33
@greptile-apps

greptile-apps Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Bundles exec-harness and memtrack into the main binary, removing separate installation.

The PR appears safe to merge; no outstanding blocking or non-blocking defect was established.

Summary

This PR consolidates the runner, execution harness, and memory tracker into one distributable binary while removing the preload-based simulation path. Changes made since the previous review additionally:

  • Repair nested-u(ret)probe stack captures before hashing them.
  • Bind memory-profile symbols and unwind information to verified mapped file bytes.
  • Add integration-test support for custom Valgrind branches, tags, and commits.
  • Expand regression coverage for nested allocator calls and overlay-backed modules.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
    CLI[codspeed binary] --> EH[Hidden exec-harness subcommand]
    CLI --> MT[Hidden memtrack subcommand]
    EH --> BENCH[Spawn benchmark]
    MT --> BPF[eBPF allocator probes]
    BPF --> STACK[Copy and repair stack bytes]
    STACK --> ART[Memtrack artifact]
    ART --> MAP[Verify mapped-file identity]
    MAP --> SYMBOLS[Extract symbols and unwind data]
Loading

Reviews (7) · Last reviewed commit: "fix(memory): grant the eBPF capabilities..." · Reviewed by Greptile

Comment thread src/executor/memory/setup.rs Outdated
Comment thread src/cli/mod.rs Outdated
Comment thread crates/exec-harness/build.rs Outdated
Comment thread .github/workflows/ci.yml
moha-bekh added a commit that referenced this pull request Sep 17, 2026
`self_exe()` resolves the binary that internal subcommands are re-invoked
through, and the memory executor hands that same path to
`sudo setcap <caps>+ep` so the capabilities land on the binary that is
actually exec'd. Reading an environment variable there means anyone able to
set one variable chooses which file receives CAP_SYS_ADMIN and CAP_BPF.

The override exists for the tests, where `current_exe()` is the test harness
and cannot dispatch a subcommand. Nothing in production sets it -- the doc
comment justified it with a launcher scenario that has no caller. Putting it
behind `cfg(test)`, constant included, removes the escalation path outright
while keeping the tests working; a release build now always resolves
`current_exe()`.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
moha-bekh added a commit that referenced this pull request Sep 17, 2026
`codspeed memtrack` and `codspeed exec-harness` are a re-exec of this binary
and share nothing with the runner, but they were dispatched at the bottom of
`run()` -- after the profile config is loaded, after the API client is built,
and after `DiscoveredProjectConfig::discover_and_load` walks the filesystem.

That last one is the problem: the re-exec runs in the benchmark's working
directory, which is the user's project. A malformed `codspeed.yaml` there
aborts the subcommand, so a measurement fails for a reason that has nothing
to do with the measurement, and a `--config` given to the outer run is not
forwarded to the inner one to override it.

Move them into `run_internal`, called right after `Cli::parse()`. The logger
match loses its internal arms for the same reason it had them.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
moha-bekh added a commit that referenced this pull request Sep 17, 2026
The released Linux artifacts are `aarch64-unknown-linux-musl` and
`x86_64-unknown-linux-musl`, and nothing in CI built either: a break in the
argp stub, in the kernel-header paths or in the aarch64 `-lgcc` link flag
would have surfaced for the first time during a tag-triggered release. The
throwaway spike workflow used to cover this and was deleted with the spike.

Both legs build on a native runner, with no environment variables, which is
also what keeps `.cargo/config.toml` honest -- it has to carry the whole
recipe on its own. The assertions are `readelf`-based rather than a `file`
string, since rustc emits a static-PIE for x86_64 musl and spells it
differently from aarch64, and `codspeed exec-harness --version` /
`codspeed memtrack --version` answer only if both CLIs really are linked in.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
moha-bekh added a commit that referenced this pull request Sep 17, 2026
`self_exe()` resolves the binary that internal subcommands are re-invoked
through, and the memory executor hands that same path to
`sudo setcap <caps>+ep` so the capabilities land on the binary that is
actually exec'd. Reading an environment variable there means anyone able to
set one variable chooses which file receives CAP_SYS_ADMIN and CAP_BPF.

The override exists for the tests, where `current_exe()` is the test harness
and cannot dispatch a subcommand. Nothing in production sets it -- the doc
comment justified it with a launcher scenario that has no caller. Putting it
behind `cfg(test)`, constant included, removes the escalation path outright
while keeping the tests working; a release build now always resolves
`current_exe()`.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
moha-bekh added a commit that referenced this pull request Sep 17, 2026
`codspeed memtrack` and `codspeed exec-harness` are a re-exec of this binary
and share nothing with the runner, but they were dispatched at the bottom of
`run()` -- after the profile config is loaded, after the API client is built,
and after `DiscoveredProjectConfig::discover_and_load` walks the filesystem.

That last one is the problem: the re-exec runs in the benchmark's working
directory, which is the user's project. A malformed `codspeed.yaml` there
aborts the subcommand, so a measurement fails for a reason that has nothing
to do with the measurement, and a `--config` given to the outer run is not
forwarded to the inner one to override it.

Move them into `run_internal`, called right after `Cli::parse()`. The logger
match loses its internal arms for the same reason it had them.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
moha-bekh added a commit that referenced this pull request Sep 17, 2026
The released Linux artifacts are `aarch64-unknown-linux-musl` and
`x86_64-unknown-linux-musl`, and nothing in CI built either: a break in the
argp stub, in the kernel-header paths or in the aarch64 `-lgcc` link flag
would have surfaced for the first time during a tag-triggered release. The
throwaway spike workflow used to cover this and was deleted with the spike.

Both legs build on a native runner, with no environment variables, which is
also what keeps `.cargo/config.toml` honest -- it has to carry the whole
recipe on its own. The assertions are `readelf`-based rather than a `file`
string, since rustc emits a static-PIE for x86_64 musl and spells it
differently from aarch64, and `codspeed exec-harness --version` /
`codspeed memtrack --version` answer only if both CLIs really are linked in.

Reported by Greptile on #531.

Refs COD-3440
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@moha-bekh
moha-bekh force-pushed the spike/cod-3440-memtrack-musl branch from 3294b83 to 8e2d7e3 Compare September 17, 2026 15:34

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@not-matthias can you do a first round of review 🙏 ? Let's ignore the LD_PRELOAD removal overhead for the exec-harness simulation, I'll spec this in a dedicated ticket and we'll do it before we merge and release this.

@not-matthias not-matthias left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall pretty good, just a few comments on how to better structure the code/comments

And a few notes on how to best structure the PRs for reviewers:

  • Try to keep the changes as minimal as possible
  • Do not modify unrelated comments/code -> should be done in separate PRs for easier review
  • Remove obvious/LLM-written/bloated comments (using the deslop skill)
  • you should also review the PR before putting it into review

Comment thread crates/exec-harness/Cargo.toml Outdated
Comment thread crates/memtrack/Cargo.toml Outdated
Comment thread CONTRIBUTING.md Outdated
Comment thread crates/exec-harness/src/analysis/mod.rs Outdated
Comment thread crates/exec-harness/src/analysis/mod.rs Outdated
Comment thread src/executor/config.rs
Comment thread crates/memtrack/src/ebpf/memtrack/mod.rs Outdated
Comment thread src/executor/tests.rs Outdated
Comment thread src/executor/memory/setup.rs Outdated
Comment thread src/run_environment/local/provider.rs

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comments for details.

I also see that there are unresolved comments.

When you handle a review, all comments should either be

  • resolved and handled properly (very important to not silently resolve comments)
  • answered with either counter arguments or questions so the reviewer can either accept your arguments or open a discussion

Finally, there are 35 commits in the PR, you may want to re-do commits before the next round of reviews

Comment thread crates/exec-harness/src/analysis/mod.rs Outdated
Comment thread src/executor/valgrind/measure.rs Outdated
Comment thread src/executor/valgrind/measure.rs
Comment thread src/executor/valgrind/measure.rs Outdated
Comment thread crates/exec-harness/build.rs Outdated
Comment thread src/cli/memtrack.rs Outdated
Comment thread src/cli/mod.rs
Comment thread src/executor/config.rs Outdated
Comment thread .cargo/config.toml Outdated
Comment thread .github/workflows/ci.yml Outdated

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comments for details.

I also see that there are unresolved comments.

When you handle a review, all comments should either be

  • resolved and handled properly (very important to not silently resolve comments)
  • answered with either counter arguments or questions so the reviewer can either accept your arguments or open a discussion

Finally, there are 35 commits in the PR, you may want to re-do commits before the next round of reviews

Comment thread .github/workflows/ci.yml Outdated

Copy link
Copy Markdown
Member Author

Yeah I didn't see them since I needed to click on the see more button, I will check If they are already resolved, and try to reduce the amount of commits by squashing where it's coherent

@moha-bekh
moha-bekh force-pushed the spike/cod-3440-memtrack-musl branch from 96d9261 to ac9a23c Compare September 21, 2026 13:05

@not-matthias not-matthias left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we're getting there!

Comment thread src/executor/memory/executor.rs Outdated
Comment thread src/executor/memory/setup.rs Outdated
Comment thread src/executor/memory/setup.rs Outdated

@not-matthias not-matthias left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we're getting there!

Comment thread src/executor/config.rs
Comment thread src/executor/memory/setup.rs Outdated
Comment thread Cargo.toml Outdated
Comment thread crates/exec-harness/src/cli.rs Outdated
Comment on lines +9 to +12
#[command(
version,
about = "CodSpeed exec harness - wraps commands with performance instrumentation"
)]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no need for this IMO, the command is hidden, and its version will simply follow the codspeed version, which we should actually update in tehc argo.toml. Same for memtrack actually.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed it, but I'm not sure I understand what I should update, Pass the crates to the codspeed version?

Comment thread src/cli/memtrack.rs Outdated
Comment thread .github/workflows/ci.yml Outdated
Comment thread src/executor/config.rs
Comment thread src/executor/orchestrator.rs Outdated
@GuillaumeLagrange
GuillaumeLagrange removed the request for review from not-matthias October 5, 2026 15:46
@moha-bekh
moha-bekh force-pushed the spike/cod-3440-memtrack-musl branch 2 times, most recently from 8a97c67 to 2b4a5c0 Compare October 6, 2026 10:23
Comment thread src/executor/memory/setup.rs Outdated
Comment thread src/executor/memory/setup.rs
Comment thread src/executor/memory/setup.rs Outdated
@moha-bekh
moha-bekh force-pushed the spike/cod-3440-memtrack-musl branch from 73f7a63 to b933455 Compare October 7, 2026 08:43
@greptile-apps

greptile-apps Bot commented Oct 7, 2026

Copy link
Copy Markdown

Want your agent to iterate on Greptile's feedback? Start a greploop in Claude Code and it will work through the open comments and keep going until this PR reviews clean.

@lvaroqui
lvaroqui added this pull request to stack #572 October 7, 2026 09:22
moha-bekh and others added 6 commits October 7, 2026 11:41
exec-harness injected a shared library into every benchmark to drive valgrind's
instrumentation from inside the child. That only works on a dynamically linked
executable, so statically linked benchmarks were silently unmeasurable, and it
forced the harness to ship a `.so` next to its binary.

The instrumentation is now toggled in exec-harness's own process, around the
spawn: valgrind propagates the state across `fork`/`exec`, so the child is
measured without anything being injected into it. The preload library, its
compatibility check and the build script that produced it all go away, and the
integration constants become plain consts.

Since the child no longer switches instrumentation on for itself, exec-harness
runs pass `--instr-atstart=inherit` to valgrind; with `no`, the child starts
uninstrumented and the measurement comes back empty. Entrypoint runs keep the
previous default.

BREAKING CHANGE: exec-harness no longer ships a preload library.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both binaries kept their argument parsing and dispatch in `main.rs`, where
nothing else can reach it. Move each into a `cli` module of its own crate and
leave `main.rs` as a wrapper that installs a logger and calls `run_cli`.

Nothing changes for the standalone binaries, but the runner can now link either
CLI and dispatch it in-process.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The runner downloaded `exec-harness` and `memtrack` from GitHub releases at the
start of a run, pinned by version, and `setup --mode memory` installed memtrack
with `cargo install`. That is a network round-trip on every run, a version
matrix to keep in sync, and two more artifacts to release.

Link both crates instead and expose them as hidden `codspeed exec-harness` and
`codspeed memtrack` subcommands, re-executing the current binary where the
runner used to invoke the downloaded tool. They are dispatched before any runner
setup: the re-exec happens in the benchmark's working directory, where an
unrelated `codspeed.yaml` would otherwise abort the measurement.

The binary installer, the memory setup no-op and the memtrack tool status go
away.

BREAKING CHANGE: `exec-harness` and `memtrack` are no longer downloaded or
installed separately; the runner binary carries them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`exec-harness` and `memtrack` were released as their own artifacts. Nothing
downloads them any more, so they stop being release units: their apt build
dependencies move to the runner, which now builds the vendored libbpf and
elfutils, and memtrack is depended on with its default features so the bundled
subcommand carries the tracker.

The released Linux artifacts are musl, and memtrack could not be built for them:
`libbpf-sys` vendors elfutils, whose `configure` looks for `argp`, `obstack` and
`fts`, none of which musl ships, and Debian's `musl-gcc` runs with `-nostdinc`,
so the kernel UAPI headers libbpf needs are out of reach. The recipe lives in
the cargo config so a plain `cargo build --target <arch>-unknown-linux-musl`
works: seed the autoconf cache for the three checks, add a declarations-only
`argp.h` stub on `CPATH`, add the UAPI header paths back through the per-target
`CFLAGS`, and link `-lgcc` on aarch64 for libbpf's outline-atomic helpers.
`close_range` goes through the raw syscall, since `libc` only declares it for
glibc.

CI builds both musl targets on native runners and asserts the artifact is
static with `readelf`. The bundled subcommands no longer take `--version`, and the
memtrack benchmarks call memtrack through the runner instead of installing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
memtrack runs as a subcommand of this binary, so the allocator set in its own
`main.rs` no longer applies to it. Set mimalloc on the runner binary instead,
which covers memtrack and the rest of the runner alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With memtrack bundled, `setcap` targeted the runner binary itself, so every
`codspeed` invocation ran in glibc's secure-execution mode: `LD_*`, `TMPDIR`
and similar variables were stripped before the runner could forward them to
the benchmark, in every mode.

Grant the capabilities to a copy of the binary in `~/.cache`, keyed by its ELF
build id, and run `codspeed memtrack` from it. The copy is installed in one
sudo call, since `~/.cache` can be root-owned, and copies of other builds are
pruned once they are a day old. The runner keeps the user's environment and no
longer holds the eBPF capabilities itself.

The memtrack benchmarks call `codspeed memtrack` from PATH, so CI grants that
binary its capabilities directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lvaroqui
lvaroqui force-pushed the spike/cod-3440-memtrack-musl branch from b933455 to 3a766f1 Compare October 7, 2026 09:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants