Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
78 changes: 73 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,6 +165,67 @@ intended to be used with any harness except `harness-ractor`.
Note: The `harness-ractor` harness is automatically selected when using these
categories, so there's no need to specify `--harness` manually.

### Ractor Scenario Benchmarks

By default, `harness-ractor` spawns the worker Ractors and runs the benchmark
block inside each of them. A benchmark that calls
`run_benchmark(n, scenario: true)` uses scenario mode instead. The harness
calls the block one time per trial in the main Ractor, with the Ractor count as
its argument. The block spawns and coordinates its own Ractors. When the block
returns a proc, the harness calls the proc after the retention measurement.

Scenario mode skips Ractor count 0 and runs no warmup. For each trial, the
harness records:

* the time of the block call, which includes all work in the block;
* the retained RSS: the RSS after a full GC, minus the RSS of the process
before the first trial;
* the peak RSS: the highest RSS that the harness reads just before the block
call, every 5 ms during it (`RACTOR_MEM_PEAK_SAMPLE_INTERVAL`), and just
after it.

Three benchmarks use scenario mode to measure pathological memory behaviour
with multiple ractors, for GC work that reclaims ractor-local memory:

* **`ractor-dead-set`** - Every ractor builds a large live set and terminates.
Retention shows how much of the dead ractors' final live sets a full GC
leaves resident.
* **`ractor-idle-garbage`** - Every ractor builds a large set, drops all
references, then idles without allocating. The garbage cannot be swept
while the ractor idles.
* **`ractor-msg-backlog`** - Unshareable payloads flood the queues of gated
consumer ractors, duplicating the payload data per consumer. Its time
includes the gate sleep (`RACTOR_BACKLOG_GATE_SLEEP`, default 1 second).

```bash
ruby -Iharness-ractor benchmarks/ractor-dead-set/benchmark.rb
```

The harness prints `BENCH_METRIC retained_mib=<worst count median>` and
`BENCH_METRIC peak_mib=...` lines, plus one pair per ractor count. The JSON
fields `ractor_mem_medians` and `ractor_mem_samples` hold the same data. The
summary table of `run_benchmarks.rb` does not show it. The ractor counts and
trials are controlled with `RUBY_BENCH_RACTORS` (default `1,2,4,6,8`) and
`MIN_BENCH_ITRS` (default 3 for these benchmarks).

The harness collects with `GC.start(global: true)` when the target Ruby's
`GC.start` accepts the `global:` keyword. Some Ruby 4.1 builds do not accept it.
On a target with Ractor-local GC, a plain `GC.start` collects only the main
Ractor's object space. The JSON field `ractor_mem_settle` records `global` or
`default`.

With `--ractor-gc` (`RUBY_BENCH_RACTOR_GC=1`), the harness cannot see which
Ractors are workers. A scenario wraps each worker body in
`measure_worker_gc { ... }`, which returns `[result, sample]`. The main Ractor
passes each sample to `record_worker_gc(worker_index, sample)`. A trial fails
when its recorded worker indexes are not `0...count`.

Worker samples cover only the workers' own object spaces during the scenario.
They do not include allocation by the main Ractor, such as the payloads that
`ractor-msg-backlog` sends. They also do not include the GCs that the harness
runs to measure retention. The JSON field `gc_controller_samples` covers the
main Ractor during the scenario.

## Ruby options

By default, ruby-bench benchmarks the Ruby used for `run_benchmarks.rb`.
Expand Down Expand Up @@ -303,7 +364,8 @@ process's lifetime peak from `getrusage`.
## Measuring Ractor GC activity

The `--ractor-gc` option of `run_benchmarks.rb` collects Ractor-local GC
metrics for benchmarks that use the Ractor harness (`--category ractor`).
metrics for benchmarks that use the Ractor harness (`--category ractor`),
in both the per-worker mode and scenario mode.
The target must use Ruby 4.1 or newer with per-Ractor global GC attribution
([ruby/ruby#19147](https://github.com/ruby/ruby/pull/19147)); older targets
fail before warmup.
Expand All @@ -317,15 +379,21 @@ Ractor's own object space. The JSON output records the scope as
`gc_scope: "ractor-local-workload"`, `gc_stat_scope: "ractor-local"`, and
`gc_measure_total_time_scope: "ractor-local"`, plus the target's `gc_config`.

The summary table adds these columns:
The text summary shows GC data in separate tables after the timing table.
A single-executable report has one `GC summary` table. A comparison report
has a `GC time ratios` table (base/comparison) and a `GC counts` table
(base → comparison). A table hides a column that has no data in any row and
lists the hidden columns below the table. A ratio column has no data when it
is `N/A` in every row; a `0.000` ratio stays visible. Any other column has no
data when it is zero or `N/A` in every row.

* `(worker sum)` columns add the Ractor-local counters of the sampled
workers of each iteration. `GCs/iter` is the sum of `minor/iter`,
* Tables marked `worker sum` add the Ractor-local counters and GC times of
the sampled workers of each iteration. `GCs/iter` is the sum of `minor/iter`,
`major/iter`, and `global/iter`; a global cycle counts under `global` on
the Ractor that initiated it, not under `major`. Single-executable reports
also show `GC ms/worker`, which divides each iteration's worker-sum GC
time by its sampled worker count, then averages.
* `controller compacts/iter*` shows the main Ractor's
* `compacts*` shows the main Ractor's
`GC.stat(:compact_count)` delta. Every global compacting cycle increments
it in every object space, so it is not summed across workers.

Expand Down
15 changes: 15 additions & 0 deletions benchmarks.yml
Original file line number Diff line number Diff line change
Expand Up @@ -296,6 +296,21 @@ json_parse_string:
ractor: true
ractor_only: true
default_harness: harness-ractor
ractor-dead-set:
desc: multiple ractors each build a large live set and terminate, then memory retention after their death is measured.
ractor: true
ractor_only: true
default_harness: harness-ractor
ractor-idle-garbage:
desc: multiple ractors each build and drop a large garbage set, then idle uncollected while retention is measured.
ractor: true
ractor_only: true
default_harness: harness-ractor
ractor-msg-backlog:
desc: unshareable message payloads flood the queues of gated consumer ractors, duplicating data per consumer and driving peak RSS and retention.
ractor: true
ractor_only: true
default_harness: harness-ractor
symbol-name-ractor:
desc: repeatedly calls Symbol#name on a static symbol under the ractor harness to stress ID-to-string lookup.
ractor: true
Expand Down
34 changes: 34 additions & 0 deletions benchmarks/ractor-dead-set/benchmark.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Multiple ractors each build a large live set and then terminate.
# The final live sets of dead ractors are garbage after death. Measured
# retention shows how much of that memory a full GC fails to reclaim.

Warning[:experimental] = false

require_relative "../../harness/loader"

DEAD_SET_ITEMS = Integer(ENV.fetch("RACTOR_DEAD_SET_ITEMS", 100_000))

run_benchmark(3, scenario: true) do |num_ractors|
workers = num_ractors.times.map do |worker|
Ractor.new(worker, DEAD_SET_ITEMS) do |worker_id, items|
measure_worker_gc do
keep = []
i = 0
while i < items
keep << "worker #{worker_id} item #{i} " + ("y" * 100)
i += 1
end
keep.size
end
end
end

total = 0
workers.each_with_index do |worker, worker_id|
size, sample = worker.value
record_worker_gc(worker_id, sample)
total += size
end
raise "unexpected dead-set size" unless total == num_ractors * DEAD_SET_ITEMS
nil
end
46 changes: 46 additions & 0 deletions benchmarks/ractor-idle-garbage/benchmark.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Multiple ractors each build a large live set, drop all references, then park
# on Ractor.receive without allocating again. Their garbage cannot be swept by
# the main Ractor's GC while they idle, so it is retained until each worker
# is released. The scenario returns a cleanup proc that the harness calls
# after the retention measurement.

Warning[:experimental] = false

require_relative "../../harness/loader"

IDLE_GARBAGE_ITEMS = Integer(ENV.fetch("RACTOR_IDLE_GARBAGE_ITEMS", 100_000))

run_benchmark(3, scenario: true) do |num_ractors|
drained = Ractor.new(num_ractors) do |count|
Array.new(count) { Ractor.receive }
end

workers = num_ractors.times.map do |worker|
Ractor.new(drained, worker, IDLE_GARBAGE_ITEMS) do |ack, worker_id, items|
keep = []
_, sample = measure_worker_gc do
i = 0
while i < items
keep << "worker #{worker_id} garbage #{i} " + ("y" * 100)
i += 1
end
end
message = Ractor.make_shareable([worker_id, sample])
keep = nil
ack.send message
Ractor.receive
:worker_done
end
end

ractor, drained_workers = Ractor.select(drained, *workers)
raise "unexpected drain barrier result" unless ractor.equal?(drained) && drained_workers.size == num_ractors
drained_workers.each { |worker_id, sample| record_worker_gc(worker_id, sample) }

proc do
workers.each do |worker|
worker.send :stop
raise "unexpected worker result" unless worker.value == :worker_done
end
end
end
47 changes: 47 additions & 0 deletions benchmarks/ractor-msg-backlog/benchmark.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# The main Ractor floods the incoming queues of multiple gated consumer
# ractors with unshareable string payloads. The copies made on send pile up
# in the queues while the consumers sleep, which duplicates the payload data
# per consumer and drives peak RSS. Measured retention after the queues are
# drained shows how much of the copied memory a full GC fails to reclaim.

Warning[:experimental] = false

require_relative "../../harness/loader"

BACKLOG_MESSAGES = Integer(ENV.fetch("RACTOR_BACKLOG_MESSAGES", 20_000))
BACKLOG_GATE_SLEEP = Float(ENV.fetch("RACTOR_BACKLOG_GATE_SLEEP", 1.0))

run_benchmark(3, scenario: true) do |num_ractors|
consumers = num_ractors.times.map do |consumer|
Ractor.new(consumer, BACKLOG_GATE_SLEEP) do |consumer_id, gate|
measure_worker_gc do
sleep gate
taken = 0
loop do
message = Ractor.receive
break if message == :done
taken += 1
end
taken
end
end
end

consumers.each_with_index do |consumer, consumer_id|
i = 0
while i < BACKLOG_MESSAGES
consumer.send "consumer #{consumer_id} message #{i} " + ("x" * 200)
i += 1
end
consumer.send :done
end

total = 0
consumers.each_with_index do |consumer, consumer_id|
taken, sample = consumer.value
record_worker_gc(consumer_id, sample)
total += taken
end
raise "unexpected backlog drain" unless total == num_ractors * BACKLOG_MESSAGES
nil
end
Loading
Loading