Skip to content

Layered/incremental snapshots - #1902

Open
ludfjig wants to merge 6 commits into
hyperlight-dev:mainfrom
ludfjig:ls-v5
Open

ludfjig wants to merge 6 commits into
hyperlight-dev:mainfrom
ludfjig:ls-v5

Conversation

@ludfjig

@ludfjig ludfjig commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Snapshots share immutable memory layers. A capture copies only the pages changed since its parent snapshot, and Sandbox::snapshot switches the sandbox to the capture so the next one is incremental too. This makes snapshot(), restore and sandbox creation much faster for larger sandboxes, and lets related snapshots share memory. See HIP 0003.

Deviations from the merged HIP 0003 draft (the HIP is updated):

  • Each snapshot stores its page tables in a separate block. A restore uses only the newest page tables, so copies in shared blobs would waste memory and disk. The guest and CPU write to page tables, so a restore must copy them into writable scratch, and keeping them in a read-only blob saves no copy.
  • Sandbox::snapshot installs the capture, because otherwise every later capture copies all pages changed since the last restore again. Installing removes map_region and map_file_cow regions, because the snapshot now holds their mapped pages and later blobs may reuse their addresses.
  • Snapshots from earlier versions are rejected instead of loaded as one layer. This is simpler, and there is little reason to keep backwards compatibility here.

Changes

  • Snapshot memory is a list of immutable shared blobs with per-snapshot live ranges, plus separate page tables copied into scratch on restore.
  • Capture reuses parent pages the guest did not change and copies the rest into one new blob.
  • Sandbox::snapshot installs the capture into the running sandbox. Scratch pages it frees are zeroed.
  • The OCI format stores ordered data layers, a page-table layer and a transport layer (config v4, ABI 6).
  • Snapshot performance: adjacent page-table walk mappings are merged, and page-table pages are cached during capture.
  • KVM: KVM_X86_QUIRK_SLOT_ZAP_ALL is disabled where supported, so deleting a memory slot flushes only that slot.
  • New benchmarks snapshots/call_and_create and snapshots/restore_alternate cover the cost moved into the next call and restores to a non-current snapshot.

Breaking changes

  • Sandbox::snapshot removes map_region and map_file_cow regions if guest has not mapped them in. Pages of these regions the guest has not mapped are lost.
  • PtRootFinder receives a SnapshotMemoryReader and returns Result<Vec<u64>>.
  • A snapshot fails if it needs more than 30 memory mappings. This bounds restore cost by to limiting the number of hypercalls to map memory.

Commits

  1. Add call_and_create and restore_alternate snapshot benchmarks. Lands first so the benchmarks run before and after.
  2. Implement layered and incremental snapshots.
  3. Install captured snapshots into the sandbox.
  4. vmem: merge adjacent mappings in walk_va_spaces.
  5. Cache page-table pages in PageTableReader.
  6. Disable KVM_X86_QUIRK_SLOT_ZAP_ALL.

Benchmarks

  • MSHV: AMD EPYC 7763, Linux 6.6 mshv.
  • KVM AMD: AMD EPYC 7763, Linux 6.17.
  • KVM Intel: Xeon Platinum 8370C, Linux 6.17.

Highlights

Benchmark MSHV KVM AMD KVM Intel
snapshots/create/default 452 µs → 241 µs (−46%) 370 µs → 283 µs (−23%) 350 µs → 195 µs (−42%)
snapshots/create/large 165 ms → 1.92 ms (−99%) 200 ms → 1.82 ms (−99%) 155 ms → 1.99 ms (−99%)
snapshots/call_and_create/medium 38.9 ms → 755 µs (−98%) 35 ms → 817 µs (−98%) 37.4 ms → 665 µs (−98%)
snapshots/restore/large 4.47 ms → 125 µs (−97%) 1.69 ms → 46.3 µs (−97%) 1.13 ms → 60 µs (−95%)
sandboxes/create_initialized/large 38.6 ms → 10.8 ms (−72%) 26.1 ms → 3.34 ms (−87%) 58.9 ms → 2.58 ms (−96%)
snapshot_files/cold_start_via_evolve/large 38.9 ms → 10.9 ms (−72%) 27.1 ms → 3.54 ms (−87%) 63.2 ms → 2.55 ms (−96%)

Guest calls and restores to the current snapshot show no consistent change.

Regressions

Benchmark MSHV KVM AMD KVM Intel Cause
snapshot_files/load_snapshot_unverified/default 87.6 µs → 134 µs (+55%) 93.4 µs → 140 µs (+50%) 48.8 µs → 70.1 µs (+44%) Maps 3 blob files instead of 1 (2 layers + page tables)
snapshot_files/load_snapshot/default 603 µs → 690 µs (+14%) 600 µs → 699 µs (+17%) 626 µs → 700 µs (+12%) Same 3 files, plus hashing 44 KiB of overwritten pages kept in the base layer
snapshot_files/save_snapshot/default 811 µs → 884 µs (+9%) 28.5 ms → 33 ms (+16%) 27.7 ms → 35.2 ms (+20%) Writes and hashes 3 blob files instead of 1
sandboxes/sandbox_from_snapshot/large 6.05 ms → 6.09 ms (~) 1.02 ms → 1.58 ms (+49%) 696 µs → 1.16 ms (+74%) KVM: 5 memory slots instead of 1, one per live range
snapshots/restore_alternate/default 184 µs → 138 µs (−24%) 235 µs → 389 µs (+66%) 146 µs → 276 µs (+89%) KVM: one memory slot change per range that differs between the two snapshots
snapshot_files/save_snapshot/small 6.24 ms → 6.33 ms (~) 36.9 ms → 40.9 ms (+11%) 37.2 ms → 42.5 ms (+14%) Writes and hashes 3 blob files instead of 1
sandboxes/sandbox_from_snapshot/default 601 µs → 596 µs (~) 765 µs → 916 µs (+15%) 520 µs → 662 µs (+23%) KVM: one memory slot per live range
snapshots/restore_alternate/small 356 µs → 137 µs (−62%) 265 µs → 431 µs (+59%) 175 µs → 295 µs (+78%) KVM: one memory slot change per range that differs between the two snapshots
snapshot_files/cold_start_via_snapshot_unverified/small 2.86 ms → 1.49 ms (−49%) 1.76 ms → 2.24 ms (+29%) 1.34 ms → 1.49 ms (+16%) KVM: the load and sandbox_from_snapshot costs combined

KVM sandbox_from_snapshot is slower at every size (+0.14 to +0.56 ms) and restore_alternate up to medium. KVM numbers are from Linux 6.17. Kernels before 6.12 cannot disable KVM_X86_QUIRK_SLOT_ZAP_ALL, so each memory slot change flushes all guest mappings. With the quirk enabled, the first call after snapshot() measured about +650 µs instead of +60 µs. WHP and aarch64 were not benchmarked. A reused layer keeps the pages its descendants overwrote, so a single snapshot can be slightly larger on disk (44 KiB for default).

Full results

All benchmarks on all three hosts.

Full results (click to expand)

Before → after (change). Negative is faster. Bold is a regression. ~ is not significant or under ±3%. KVM AMD compares against main for benchmarks that existed before this PR.

snapshots

Benchmark Size MSHV KVM AMD KVM Intel
call_and_create default 576 µs → 367 µs (−35%) 453 µs → 397 µs (−12%) 407 µs → 296 µs (−26%)
call_and_create small 1.81 ms → 415 µs (−77%) 2.37 ms → 453 µs (−81%) 2.44 ms → 321 µs (−86%)
call_and_create medium 38.9 ms → 755 µs (−98%) 35 ms → 817 µs (−98%) 37.4 ms → 665 µs (−98%)
call_and_create large 164 ms → 2.05 ms (−99%) 146 ms → 2.04 ms (−98%) 143 ms → 1.79 ms (−99%)
create default 452 µs → 241 µs (−46%) 370 µs → 283 µs (−23%) 350 µs → 195 µs (−42%)
create small 1.64 ms → 297 µs (−82%) 1.96 ms → 339 µs (−83%) 2.31 ms → 253 µs (−89%)
create medium 38.7 ms → 659 µs (−98%) 34.5 ms → 685 µs (−98%) 36.9 ms → 554 µs (−98%)
create large 165 ms → 1.92 ms (−99%) 200 ms → 1.82 ms (−99%) 155 ms → 1.99 ms (−99%)
restore default 89.7 µs → 88.7 µs (~) 19.7 µs → 20.1 µs (~) 13.1 µs → 13.1 µs (~)
restore small 92.9 µs → 89.4 µs (−7%) 20.9 µs → 20.9 µs (~) 13.6 µs → 13.7 µs (~)
restore medium 201 µs → 95.4 µs (−69%) 52.8 µs → 24.7 µs (−69%) 55.8 µs → 17.2 µs (−82%)
restore large 4.47 ms → 125 µs (−97%) 1.69 ms → 46.3 µs (−97%) 1.13 ms → 60 µs (−95%)
restore_alternate default 184 µs → 138 µs (−24%) 235 µs → 389 µs (+66%) 146 µs → 276 µs (+89%)
restore_alternate small 356 µs → 137 µs (−62%) 265 µs → 431 µs (+59%) 175 µs → 295 µs (+78%)
restore_alternate medium 1.21 ms → 147 µs (−88%) 249 µs → 404 µs (+60%) 280 µs → 255 µs (~)
restore_alternate large 4.38 ms → 182 µs (−96%) 1.02 ms → 423 µs (−59%) 1.15 ms → 365 µs (−69%)

guest_calls

Benchmark Size MSHV KVM AMD KVM Intel
call default 44.8 µs → 41.5 µs (~) 17.8 µs → 17.3 µs (~) 17.2 µs → 18.2 µs (+7%)
call small 45.2 µs → 44.2 µs (~) 17.9 µs → 17.9 µs (~) 18.2 µs → 17.8 µs (−4%)
call medium 44.7 µs → 42.6 µs (−4%) 17.5 µs → 17.9 µs (~) 18 µs → 17.2 µs (~)
call large 44.4 µs → 43.3 µs (~) 17.6 µs → 18.2 µs (~) 18.1 µs → 18.4 µs (~)
call_with_host_function default 73.9 µs → 73.9 µs (~) 34.1 µs → 34.2 µs (~) 34.3 µs → 31.8 µs (−7%)
call_with_host_function small 76.4 µs → 75 µs (~) 34 µs → 34 µs (~) 33.5 µs → 34.8 µs (+4%)
call_with_host_function medium 76 µs → 72.5 µs (−5%) 33.8 µs → 34 µs (~) 34.4 µs → 34.8 µs (~)
call_with_host_function large 74.6 µs → 74.9 µs (~) 34.2 µs → 34.4 µs (~) 34 µs → 34.6 µs (~)
call_with_restore default 159 µs → 155 µs (~) 59.2 µs → 61.2 µs (~) 38.6 µs → 37.4 µs (~)
call_with_restore small 162 µs → 161 µs (~) 62.2 µs → 63.3 µs (~) 39.1 µs → 38.8 µs (~)
call_with_restore medium 166 µs → 162 µs (−4%) 67.4 µs → 67.9 µs (~) 43.5 µs → 43.2 µs (~)
call_with_restore large 200 µs → 184 µs (−12%) 91.5 µs → 89.1 µs (−6%) 99.6 µs → 96 µs (−4%)
different_thread - 65.6 µs → 69.7 µs (+4%) 41.2 µs → 41.3 µs (~) 72.9 µs → 75.5 µs (+3%)
interrupt_latency - 45 µs → 44.2 µs (~) 22.9 µs → 23 µs (~) 39.5 µs → 17.4 µs (~)

guest_functions_with_large_parameters

Benchmark Size MSHV KVM AMD KVM Intel
guest_call_with_large_parameters - 163 ms → 174 ms (+7%) 964 ms → 853 ms (−11%) 690 ms → 624 ms (−9%)

sample_workloads

Benchmark Size MSHV KVM AMD KVM Intel
24K_in_8K_out_c - 54.4 µs → 54.2 µs (~) 25.9 µs → 27.5 µs (+4%) 24.3 µs → 24.5 µs (~)
24K_in_8K_out_rust - 55 µs → 54.4 µs (~) 27.3 µs → 27.4 µs (~) 28.2 µs → 27.9 µs (~)

sandboxes

Benchmark Size MSHV KVM AMD KVM Intel
create_initialized default 1.42 ms → 1.42 ms (~) 2.42 ms → 2.42 ms (~) 1.75 ms → 1.78 ms (~)
create_initialized small 4.22 ms → 3.07 ms (−29%) 4.45 ms → 2.61 ms (−41%) 3.52 ms → 1.87 ms (−46%)
create_initialized medium 11.3 ms → 5.47 ms (−52%) 9.05 ms → 2.83 ms (−68%) 16.4 ms → 2.04 ms (−87%)
create_initialized large 38.6 ms → 10.8 ms (−72%) 26.1 ms → 3.34 ms (−87%) 58.9 ms → 2.58 ms (−96%)
create_initialized_and_drop default 20.4 ms → 18.9 ms (~) 10 ms → 11.8 ms (~) 9.68 ms → 7.81 ms (−17%)
create_initialized_and_drop small 28.7 ms → 25.1 ms (~) 15.1 ms → 12.3 ms (~) 11.5 ms → 8.27 ms (−34%)
create_initialized_and_drop medium 30.3 ms → 30.4 ms (~) 18.9 ms → 11.2 ms (−38%) 22.5 ms → 8.72 ms (−61%)
create_initialized_and_drop large 57 ms → 29.9 ms (−47%) 34.5 ms → 17.1 ms (−54%) 66.8 ms → 9.42 ms (−86%)
sandbox_from_snapshot default 601 µs → 596 µs (~) 765 µs → 916 µs (+15%) 520 µs → 662 µs (+23%)
sandbox_from_snapshot small 792 µs → 854 µs (+9%) 822 µs → 1.07 ms (+41%) 564 µs → 784 µs (+41%)
sandbox_from_snapshot medium 2.01 ms → 2.03 ms (~) 874 µs → 1.29 ms (+40%) 572 µs → 834 µs (+49%)
sandbox_from_snapshot large 6.05 ms → 6.09 ms (~) 1.02 ms → 1.58 ms (+49%) 696 µs → 1.16 ms (+74%)

snapshot_files

Benchmark Size MSHV KVM AMD KVM Intel
cold_start_via_evolve default 1.5 ms → 1.43 ms (~) 2.84 ms → 2.61 ms (−10%) 2.03 ms → 1.84 ms (−8%)
cold_start_via_evolve small 3.41 ms → 2.95 ms (−15%) 4.73 ms → 2.72 ms (−44%) 3.82 ms → 2.06 ms (−46%)
cold_start_via_evolve medium 13 ms → 5.52 ms (−57%) 9.1 ms → 2.97 ms (−68%) 16.8 ms → 2.09 ms (−87%)
cold_start_via_evolve large 38.9 ms → 10.9 ms (−72%) 27.1 ms → 3.54 ms (−87%) 63.2 ms → 2.55 ms (−96%)
cold_start_via_snapshot default 1.9 ms → 1.8 ms (~) 2.45 ms → 2.53 ms (+7%) 1.83 ms → 1.79 ms (~)
cold_start_via_snapshot small 9.35 ms → 7.84 ms (−16%) 8.33 ms → 8.58 ms (+4%) 8.46 ms → 8.83 ms (+4%)
cold_start_via_snapshot medium 51.8 ms → 49.2 ms (−5%) 50 ms → 50.5 ms (~) 58.6 ms → 59.3 ms (~)
cold_start_via_snapshot large 194 ms → 190 ms (~) 192 ms → 193 ms (~) 229 ms → 231 ms (~)
cold_start_via_snapshot_unverified default 1.24 ms → 1.28 ms (~) 1.96 ms → 2.09 ms (+8%) 1.33 ms → 1.33 ms (~)
cold_start_via_snapshot_unverified small 2.86 ms → 1.49 ms (−49%) 1.76 ms → 2.24 ms (+29%) 1.34 ms → 1.49 ms (+16%)
cold_start_via_snapshot_unverified medium 4.21 ms → 2.65 ms (−37%) 1.91 ms → 2.34 ms (+26%) 1.44 ms → 1.36 ms (~)
cold_start_via_snapshot_unverified large 8.27 ms → 7.05 ms (−16%) 2.15 ms → 2.55 ms (+21%) 1.54 ms → 1.79 ms (+18%)
load_snapshot default 603 µs → 690 µs (+14%) 600 µs → 699 µs (+17%) 626 µs → 700 µs (+12%)
load_snapshot small 6.34 ms → 6.4 ms (~) 6.28 ms → 6.43 ms (~) 7.22 ms → 7.37 ms (~)
load_snapshot medium 46.2 ms → 46.4 ms (~) 47.8 ms → 48.7 ms (~) 57.2 ms → 65.9 ms (+15%)
load_snapshot large 183 ms → 184 ms (~) 190 ms → 190 ms (~) 239 ms → 229 ms (−4%)
load_snapshot_unverified default 87.6 µs → 134 µs (+55%) 93.4 µs → 140 µs (+50%) 48.8 µs → 70.1 µs (+44%)
load_snapshot_unverified small 88 µs → 135 µs (+53%) 94.8 µs → 141 µs (+49%) 49.1 µs → 70.8 µs (+45%)
load_snapshot_unverified medium 89.5 µs → 136 µs (+52%) 94.6 µs → 143 µs (+52%) 52.1 µs → 72.1 µs (+38%)
load_snapshot_unverified large 93.7 µs → 166 µs (+66%) 97.7 µs → 147 µs (+51%) 53.3 µs → 73.8 µs (+38%)
save_snapshot default 811 µs → 884 µs (+9%) 28.5 ms → 33 ms (+16%) 27.7 ms → 35.2 ms (+20%)
save_snapshot small 6.24 ms → 6.33 ms (~) 36.9 ms → 40.9 ms (+11%) 37.2 ms → 42.5 ms (+14%)
save_snapshot medium 45.4 ms → 44.9 ms (~) 82.5 ms → 84.4 ms (~) 84.9 ms → 93.1 ms (+10%)
save_snapshot large 179 ms → 175 ms (~) 215 ms → 217 ms (~) 253 ms → 248 ms (~)

call_and_create times a guest call plus snapshot creation, so the call pays for memory remapping done by the previous snapshot.

restore_alternate restores two snapshots in turn, so each restore targets memory that is not current.

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
@ludfjig ludfjig added regen-goldens Regenerate snapshot golden fixtures kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. labels Oct 8, 2026
Share immutable data blobs and store page tables separately. Persist layered snapshots as OCI image layouts.

Capture copies only pages that changed since the snapshot the sandbox was created from or last restored to.

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
Snapshot capture switches the running sandbox to the captured snapshot. The next capture reuses its layers and copies only pages changed after it. Scratch pages used before the capture become free again. They are zeroed so guest allocations return zeroed pages, as after restore.

Capture removes map_region and map_file_cow regions. Pages of these regions that the guest has not mapped are lost.

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
walk_va_spaces merges mappings that have contiguous physical and virtual addresses and the same kind. Snapshot capture classifies and rebuilds each merged mapping once.

Sandbox::snapshot on KVM, release build:

* large heap: 22% to 37% faster
* 1 GiB heap: 26% to 43% faster

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
A page-table walk reads every entry through GuestPhysicalMemoryView,
which resolves the backing and copies 8 bytes through a locked volatile
copy. The reader keeps a copy of the last table page it read, so each
table costs one copy. Snapshot capture and crash dumps use the reader.

The walk of a default sandbox drops from 187 us to 44 us. Criterion
snapshots/create on MSHV:

* default: 425 us to 252 us (-36%)
* medium: 977 us to 645 us (-34%)

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
By default, deleting a memslot flushes the guest mappings of every
memslot. KVM keeps that flush as a workaround for VMs with assigned
GPUs. Hyperlight assigns no devices and disables the quirk when the
kernel supports it (Linux 6.12 and later). Deleting a memslot then
flushes only that memslot's mappings.

Snapshot capture deletes a memslot when a layer drops out of the
snapshot. In a loop of one call and one snapshot on KVM (Linux 6.18,
nested under WSL2):

* call after snapshot: 914-993 us to 143-185 us
* snapshot: 595-722 us to 414-640 us

Restore to another snapshot and region unmapping also delete memslots.

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
@ludfjig
ludfjig marked this pull request as ready for review October 8, 2026 02:06
Copilot AI balanced review requested due to automatic review settings October 8, 2026 02:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes snapshot ABI and guest memory ownership across multiple hypervisor backends, including unbenchmarked paths.

2 open findings
What changed in this PR

Adds incremental snapshots that share immutable memory layers, reducing capture and restore costs.

Changes:

  • Introduces layered snapshot memory and page-table storage.
  • Installs captures for incremental reuse across snapshots.
  • Updates OCI ABI, hypervisor mappings, tests, docs, and benchmarks.
File Description
src/​tests/​rust_guests/​simpleguest/​src/​main.rs Adds zero-page mapping fixture.
src/​hyperlight_host/​tests/​snapshot_goldens/​goldens_version.rs Bumps goldens to v6.
src/​hyperlight_host/​src/​sandbox/​uninitialized.rs Updates mapping size semantics and docs.
src/​hyperlight_host/​src/​sandbox/​trace/​mem_profile.rs Reads layered snapshot GPAs.
src/​hyperlight_host/​src/​sandbox/​snapshot/​tripwires.rs Pins ABI 6 media types.
src/​hyperlight_host/​src/​sandbox/​snapshot/​memory/​mod.rs Defines layered snapshot memory.
src/​hyperlight_host/​src/​sandbox/​snapshot/​memory/​backing.rs Maps and reads snapshot layers.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​mod.rs Persists and loads layered OCI snapshots.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​media_types.rs Adds v4/v2/page-table media types.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​config.rs Defines layered config schema.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file_tests.rs Expands layered-format tests.
src/​hyperlight_host/​src/​sandbox/​mod.rs Exports snapshot memory reader.
src/​hyperlight_host/​src/​sandbox/​initialized.rs Installs captured snapshots incrementally.
src/​hyperlight_host/​src/​mem/​virtq/​tests.rs Updates scratch restore expectations.
src/​hyperlight_host/​src/​mem/​shared_mem.rs Adds immutable memory freezing and range mapping.
src/​hyperlight_host/​src/​mem/​mgr.rs Resolves layered physical memory.
src/​hyperlight_host/​src/​mem/​layout.rs Removes flat snapshot layout fields.
src/​hyperlight_host/​src/​hypervisor/​virtual_machine/​kvm/​x86_64.rs Disables KVM slot-zap quirk.
src/​hyperlight_host/​src/​hypervisor/​virtual_machine/​hvf/​mod.rs Adapts HVF auxiliary mapping.
src/​hyperlight_host/​src/​hypervisor/​hyperlight_vm/​x86_64.rs Uses layered mappings on x86-64.
src/​hyperlight_host/​src/​hypervisor/​hyperlight_vm/​test_support.rs Updates mapping fault support.
src/​hyperlight_host/​src/​hypervisor/​hyperlight_vm/​mod.rs Manages multiple snapshot slots.
src/​hyperlight_host/​src/​hypervisor/​hyperlight_vm/​aarch64.rs Uses layered mappings on AArch64.
src/​hyperlight_host/​src/​hypervisor/​gdb/​mod.rs Resolves debugger access by backing.
src/​hyperlight_host/​build.rs Removes obsolete memory cfg.
src/​hyperlight_host/​benches/​benchmarks.rs Adds incremental snapshot benchmarks.
src/​hyperlight_common/​src/​vmem.rs Coalesces adjacent walk mappings.
src/​hyperlight_common/​src/​arch/​amd64/​vmem.rs Applies mapping coalescing on AMD64.
src/​hyperlight_common/​src/​arch/​aarch64/​vmem.rs Applies mapping coalescing on AArch64.
proposals/​0003-incremental-snapshots/​README.md Updates the incremental snapshot design.
docs/​snapshot-versioning.md Documents ABI 6 versioning.
docs/​snapshot-oci-format.md Documents layered OCI storage.
docs/​paging-development-notes.md Describes incremental page capture.
CHANGELOG.md Records breaking snapshot changes.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

let quirks = vm_fd.check_extension_raw(KVM_CAP_DISABLE_QUIRKS2.into());
// `quirks` has one bit per quirk this kernel can disable.
if quirks <= 0 || quirks & KVM_X86_QUIRK_SLOT_ZAP_ALL as i32 == 0 {
tracing::warn!("KVM does not support disabling KVM_X86_QUIRK_SLOT_ZAP_ALL");
Comment on lines +26 to +27
Page tables have their own `MT_PAGE_TABLES_V1` descriptor after the
data layers.
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: c57cef939476
Baseline commit: 5ddb37af894e

kvm / amd (Linux) (❌ 1.51x | 🚀 6.11x)

Top improvements

  • sandboxes/create_initialized_and_drop/medium — 🚀 6.11x faster

Top regressions

  • snapshot_files/load_snapshot_unverified/small — ❌ 1.51x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 782.06 ns (➖ 1.02x slower)
vec_bytes 577.38 ns (➖ 1.02x faster)
373.37 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 524.81 ns (➖ 1.07x faster)
65536 162.18 ns (➖ 1.16x slower)

sandboxes

create_initialized_and_drop
medium 12.98 ms (🚀 6.11x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.24 ns (➖ 1.03x slower) 7.73 ns (➖ 1.00x faster) 7.78 ns (➖ 1.07x faster)

snapshot_files

load_snapshot_unverified
small 141.58 µs (❌ 1.51x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.30 µs (➖ 1.00x slower) 7.25 µs (➖ 1.03x faster)
65536 2.07 µs (➖ 1.01x faster) 2.06 µs (➖ 1.07x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.31 µs (➖ 1.01x slower) 6.30 µs (➖ 1.01x faster)
8192 1.06 µs (➖ 1.02x slower) 1.04 µs (➖ 1.03x faster)
262144 26.98 µs (➖ 1.00x faster)
kvm / intel (Linux) (❌ 1.71x | 🚀 4.95x)

Top improvements

  • sandboxes/create_initialized_and_drop/medium — 🚀 4.95x faster

Top regressions

  • snapshot_files/load_snapshot_unverified/small — ❌ 1.71x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 816.46 ns (➖ 1.17x slower)
vec_bytes 594.22 ns (➖ 1.15x slower)
730.80 µs (➖ 1.11x slower)

payload_allocation

slot_pool_segmented
262144 638.00 ns (➖ 1.27x slower)
65536 163.40 ns (➖ 1.20x slower)

sandboxes

create_initialized_and_drop
medium 14.49 ms (🚀 4.95x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.26 ns (➖ 1.20x slower) 8.28 ns (➖ 1.20x slower) 8.33 ns (➖ 1.20x slower)

snapshot_files

load_snapshot_unverified
small 75.27 µs (❌ 1.71x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.98 µs (➖ 1.19x slower) 9.04 µs (➖ 1.19x slower)
65536 2.53 µs (➖ 1.19x slower) 2.58 µs (➖ 1.21x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.41 µs (➖ 1.20x slower) 8.37 µs (➖ 1.17x slower)
8192 850.83 ns (➖ 1.18x slower) 839.14 ns (➖ 1.15x slower)
262144 36.22 µs (➖ 1.23x slower)
mshv3 / amd (Linux) (❌ 1.56x)

Top regressions

  • snapshot_files/load_snapshot_unverified/small — ❌ 1.56x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 919.95 ns (➖ 1.03x faster)
vec_bytes 690.45 ns (➖ 1.04x faster)
279.49 µs (➖ 1.30x faster)

payload_allocation

slot_pool_segmented
262144 723.88 ns (➖ 1.01x slower)
65536 195.47 ns (➖ 1.10x slower)

sandboxes

create_initialized_and_drop
medium 59.44 ms (➖ 1.04x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.81 ns (➖ 1.00x slower) 9.86 ns (➖ 1.00x slower) 9.67 ns (➖ 1.01x faster)

snapshot_files

load_snapshot_unverified
small 129.06 µs (❌ 1.56x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 11.52 µs (➖ 1.13x slower) 8.95 µs (➖ 1.09x faster)
65536 2.26 µs (➖ 1.07x faster) 2.49 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.15 µs (➖ 1.05x faster) 8.32 µs (➖ 1.03x slower)
8192 1.25 µs (➖ 1.04x faster) 1.33 µs (➖ 1.02x faster)
262144 36.23 µs (➖ 1.01x faster)
mshv3 / intel (Linux) (❌ 1.50x)

Top regressions

  • sandboxes/create_initialized_and_drop/medium — ❌ 1.50x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 914.21 ns (➖ 1.02x faster)
vec_bytes 658.23 ns (➖ 1.01x faster)
668.95 µs (➖ 1.04x faster)

payload_allocation

slot_pool_segmented
262144 620.55 ns (➖ 1.10x faster)
65536 164.60 ns (➖ 1.03x faster)

sandboxes

create_initialized_and_drop
medium 96.16 ms (❌ 1.50x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.60 ns (➖ 1.03x faster) 8.78 ns (➖ 1.00x slower) 8.77 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 66.23 µs (➖ 1.46x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.76 µs (➖ 1.01x faster) 7.86 µs (➖ 1.01x slower)
65536 2.27 µs (➖ 1.01x slower) 2.29 µs (➖ 1.01x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.53 µs (➖ 1.00x slower) 7.59 µs (➖ 1.01x faster)
8192 872.88 ns (➖ 1.01x slower) 854.61 ns (➖ 1.04x faster)
262144 38.58 µs (➖ 1.11x slower)
hyperv-ws2025 / amd (Windows) (🚀 6.51x)

Top improvements

  • sandboxes/create_initialized_and_drop/medium — 🚀 6.51x faster

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.16 µs (➖ 1.00x slower)
vec_bytes 771.33 ns (➖ 1.17x faster)
2.14 ms (➖ 1.03x faster)

payload_allocation

slot_pool_segmented
262144 799.59 ns (➖ 1.01x faster)
65536 230.45 ns (➖ 1.00x faster)

sandboxes

create_initialized_and_drop
medium 13.03 ms (🚀 6.51x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.75 ns (➖ 1.06x slower) 10.59 ns (➖ 1.04x slower) 10.55 ns (➖ 1.06x slower)

snapshot_files

load_snapshot_unverified
small 784.50 µs (➖ 1.11x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.77 µs (➖ 1.08x faster) 9.36 µs (➖ 1.16x faster)
65536 2.55 µs (➖ 1.05x slower) 2.75 µs (➖ 1.15x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.66 µs (➖ 1.02x faster) 8.82 µs (➖ 1.03x faster)
8192 1.40 µs (➖ 1.04x slower) 1.44 µs (➖ 1.13x slower)
262144 38.46 µs (➖ 1.11x faster)
hyperv-ws2025 / intel (Windows) (❌ 1.51x | 🚀 6.89x)

Top improvements

  • sandboxes/create_initialized_and_drop/medium — 🚀 6.89x faster

Top regressions

  • snapshot_files/load_snapshot_unverified/small — ❌ 1.51x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.33 µs (➖ 1.06x slower)
vec_bytes 822.21 ns (➖ 1.04x slower)
3.34 ms (➖ 1.06x slower)

payload_allocation

slot_pool_segmented
262144 728.01 ns (➖ 1.40x faster)
65536 227.00 ns (➖ 1.24x faster)

sandboxes

create_initialized_and_drop
medium 15.73 ms (🚀 6.89x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.98 ns (➖ 1.32x faster) 10.47 ns (➖ 1.40x faster) 9.91 ns (➖ 1.30x faster)

snapshot_files

load_snapshot_unverified
small 909.75 µs (❌ 1.51x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.67 µs (➖ 1.01x slower) 10.35 µs (➖ 1.29x slower)
65536 2.32 µs (➖ 1.01x slower) 2.95 µs (➖ 1.25x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.40 µs (➖ 1.01x faster) 9.31 µs (➖ 1.09x slower)
8192 1.14 µs (➖ 1.09x faster) 1.11 µs (➖ 1.01x faster)
262144 46.24 µs (➖ 1.09x slower)

Reported by cargo ci bench-report --candidate run:37716086663 --baseline run:37707245718 --config-file bench_report.toml.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. regen-goldens Regenerate snapshot golden fixtures

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants