Skip to content

🎟️ fix: Never Reuse MicroVM Launch Tokens After a Counter Reset - #322

Draft
TomasPalsson wants to merge 1 commit into
LibreChat-AI:mainfrom
aproorg:fix/microvm-clienttoken-reuse
Draft

TomasPalsson wants to merge 1 commit into
LibreChat-AI:mainfrom
aproorg:fix/microvm-clienttoken-reuse

Conversation

@TomasPalsson

Copy link
Copy Markdown

Problem

On the lambda-microvm backend, stateful code runs fail with HTTP 503 microvm_launch_failed when a user comes back after an idle night. The worker logs:

MICROVM_LAUNCH_FAILED: The provided clientToken was used with different request parameters.
(ValidationException, HTTP 400)

Every retry then fails the same way, quickly, until roughly a day after the user's first launch. After that the same request starts working again with no change on our side.

Root cause

  • Session launches use the clientToken sess-<runtimeSessionId>-<generation> (runtimeSessionLaunchClientToken).
  • The generation comes from rtsx:gen:<id>, which expires together with the session record after RUNTIME_SESSION_RECORD_TTL_SECONDS (max duration + 10 min).
  • When the key is missing, allocateRuntimeSessionGenerationScript returns the caller's seed. That seed is runtimeSessionLaunchGenerationSeed(config), which depends only on the launch config. So after an idle night, an unchanged deployment re-issues exactly the token the session's first launch used.
  • AWS still remembers that token and rejects the reuse, even though our RunMicrovm body is byte-identical. In production we saw:
    • In eu-west-1, a reused token was rejected 24h09m after its first use and accepted 24h12m after, returning the original microVM id.
    • eu-central-1 appears to retain tokens longer.
    • Reuses after the retention window succeeded every time, which is why this only bites next-day returns.
  • The rejection maps to a non-transient MICROVM_LAUNCH_FAILED. So retireExhaustedLaunchIntent keeps the PENDING intent, and canReplayPendingLaunch resends the same rejected token on every request. That is the run of fast 503s.

Fix

  • Salted seeds. runtimeSessionLaunchGenerationSeed and hostedAppLaunchGenerationSeed mix randomBytes(16) into the 52-bit seed, so a reset counter always starts at a new generation.
    • Token shape, numeric range and length are unchanged, so the -r1 reservation and older-worker PENDING takeover still hold.
    • The Lua allocate script only ever raises a live counter, so generations stay monotonic.
    • The odds of colliding with an earlier token of the same session are about n/2^52.
  • Retire rejected tokens. retireExhaustedLaunchIntent also retires a PENDING intent when AWS rejects its token as reused. AWS launched nothing for that request, so the next request allocates a fresh generation instead of replaying. AWS signals this only as a generic ValidationException, so the check matches the message text. The regex accepts both the AWS wording and the fake client's wording.

Tests

New tests in lambda-microvm.test.ts:

  • never reissues a launch token after registry loss on an unchanged config
  • stops replaying a PENDING token that AWS rejected as reused

Both fail on main with Expected: not "sess-rt_session_1-3972529254911613", and pass with this change.

Existing tests that used the seed as a constant fixture now compute it once. The two seed tests assert the salt, since config-sensitivity alone is now trivially true.

Local run with bun 1.3.14:

  • bun test src/sandbox-backend/lambda-microvm.test.ts src/hosted-app: 158 pass, 0 fail.
  • bun test ./src: the same failing set as main on this machine (Redis/egress tests that need a real Redis), plus the two new passing tests.

Notes

  • Hosted apps only get the salt. HostedAppControlPlane still treats a launch rejection as a transient hosted_app_launch_failed and keeps the PENDING intent, per its documented policy (control-plane.ts around the hosted_app_boot_failed branch). New poisoning is prevented; I left that policy alone.
  • Orphan risk. If AWS ever rejected an identical replay of an accepted token (lost response) as "different parameters", retiring would orphan that VM until idle or max-duration expiry. Before this change, the session was unusable for that same window.

Stateful runtime sessions launch with clientToken sess-<rt>-<generation>.
The rtsx:gen counter expires with the session record (max duration + 10
minutes), and the next allocation returned the fixed config-derived seed
again. A user returning after an idle night therefore resent the token of
their first launch. AWS still remembers it and rejects the reuse with
"The provided clientToken was used with different request parameters",
even for a byte-identical request. That validation failure is not
transient, so the PENDING intent kept the token and every later request
replayed it until AWS forgot the token (about 24h in eu-west-1).

- Salt the session and hosted-app generation seeds with random bytes so a
  reset counter always starts at a new generation. Token shape, range and
  length are unchanged; the allocate script still never lowers a live
  counter.
- Retire a PENDING intent when AWS rejects its token as reused. AWS
  launched nothing, so the next request allocates a fresh generation.
- Tests: registry loss on an unchanged config must not resend the token,
  and a reuse rejection must not be replayed. Fixtures that used the seed
  as a constant compute it once; seed tests assert the salt.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant