Skip to content

test(expo): add a verify skill that drives the expo-native fixture on a local device - #10087

Draft
mikepitre wants to merge 7 commits into
mike/expo-verify-hostfrom
mike/expo-verify-remote
Draft

mikepitre wants to merge 7 commits into
mike/expo-verify-hostfrom
mike/expo-verify-remote

Conversation

@mikepitre

@mikepitre mikepitre commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds verify-clerk-expo, a skill and a CLI that an agent or a developer uses to prove a @clerk/expo change on an iOS simulator or an Android emulator. It builds the expo-native fixture, launches it with the inputs from #10052, runs tests against a Clerk application that it creates and deletes, and keeps a video and screenshots of each run. The device is on the machine that runs the CLI.

It sits on #10052. #10090, on top of this, runs the same tests in CI. #10131, on top of #10090, lets a machine that cannot run a device borrow one on a CI runner.

The seven commits are in dependency order, and each adds one part.

Commit Lines What it does
1. Package, entry point, repo wiring 75 and 1,783 of generated package-lock.json Sets the skill up as a small Node package: its dependencies, the control-clerk-expo command, and the ignore and lint settings that keep its files out of the rest of the repo's way. Most of the lines are the generated lockfile.
2. Shared core 4,068 The tool itself. It gives each worktree its own simulator or emulator so two agents never drive the same one, starts the app with the inputs a test asks for, signs test users in without showing any key to the tests, and keeps the video, screenshots, and log of every run.
3. Throwaway Clerk application 1,813 Creates a real Clerk application for the session to test against, sets it up the same way every time, and deletes it with all its users when the session ends.
4. Platform host and the local devices 2,218 The part that knows this platform: how to build the expo-native fixture, install it on the simulator or emulator, and launch it.
5. Golden tests, fixtures, feature files 1,087 The tests themselves, grouped by feature, and one short page per feature that says how a user reaches it and what on screen proves it works.
6. Unit tests and their CI job 6,280 Tests of the tool's own code. They run with no simulator or emulator, no network, and no key, and a CI job runs them.
7. Docs 429 The instructions an agent reads to use the tool, and the notes that point to it from the repo's contributor docs.
Commits 2 and 3, specs/fixtures.ts in commit 5, and most of commit 6 are the same files, byte for byte, in clerk/clerk-ios#629 and clerk/clerk-android#1046. src/core/MANIFEST lists them with their hashes, and a unit test fails when one drifts. Review them once, in whichever of the three pull requests you read first.

The unit tests for all of the code arrive together in commit 6, so the commits before it do not pass a test run by themselves.

The skill is in .claude/skills/verify-clerk-expo/, and .cursor/skills/verify-clerk-expo is a symlink to it. The CLI is bin/control-clerk-expo. packages/expo/AGENTS.md has the short instructions and the root AGENTS.md points to them.

To review, read SKILL.md, then src/host.ts and src/fixture.ts, then specs/golden/. Most of the 17,753 added lines are the lockfile and files shared with the verification skills of clerk-ios (clerk/clerk-ios#629) and clerk-android (clerk/clerk-android#1046): src/core/, src/platform/ios/, src/platform/android/, specs/fixtures.ts, testing/, and every test file but test/host.test.ts and test/freshness.test.ts. src/core/MANIFEST lists the hashes of the core files, and a unit test and doctor fail when one differs. .prettierignore leaves the shared files alone.

To try it on a Mac with Xcode, Node 24.8 or newer, and the team's Clerk Platform API key:

$ pnpm install
$ npm ci --prefix .claude/skills/verify-clerk-expo
$ .claude/skills/verify-clerk-expo/bin/control-clerk-expo doctor --platform ios
$ .claude/skills/verify-clerk-expo/bin/control-clerk-expo run native-auth-view --platform ios
$ .claude/skills/verify-clerk-expo/bin/control-clerk-expo down

doctor only reads, and prints a fix for each thing the machine lacks. run builds the fixture as a Debug dev client, creates the application, takes a simulator that the CLI cloned for itself, starts Metro and tsdown --watch in packages/expo, and runs the tests. A later JS edit reaches the app on the next run with no native build. down releases the device and deletes the application with every user in it. The evidence stays in .verify/runs/<run-id>/.

The 22 tests are in specs/golden/, in seven groups.

Group What its tests cover
native-auth-view Opening and closing AuthView, its React Native logo, and a sign-in through it
user-button-and-profile The UserButton, the profile it opens, the home's sign-out, and an inline UserProfileView with onHostBack and a custom page
custom-flow-sign-in An email code sign-in on useSignIn
custom-flow-sign-up An email and password sign-up on useSignUp
token-cache-persistence The same session after a relaunch, and a signed-out start with new storage
native-js-sync A native sign-in and a native sign-out reaching useAuth, useUser, and useSession
native-modules useSignInWithGoogle opening the native Google sign-in and reporting a cancel, and useBiometricCredentials giving the native module's answer

Every test starts at the fixture's home and taps to the screen it needs, as a user would. host.launch takes no screen. Its options still choose who is signed in, the mode of AuthView, and whether storage is kept from the last launch. specs/native.ts holds the locators of the home's buttons, and a unit test fails when they differ from the fixture's.

The tests assert on what a user sees. They look in the prebuilt views first, for example the address on the code screen of AuthView or the email under Manage account. For the outcome of a flow they read the fixture's home, which shows Signed out and Sign in, or Signed in as <email>, the user ID, the session ID, and Sign out. Android tests find the prebuilt views by text, because the clerk-android release that @clerk/expo pins has no test tags. iOS tests use the clerk.* accessibility identifiers. Tests type only +clerk_test addresses and the test code 424242. Seven tests type a code or a password and carry the tag form-entry, so a runtime that must not type them passes --skip form-entry. Three tests are for iOS only, so Android runs 19.

The two native-modules tests sign no one in. The Google test taps Sign in with Google and cancels what opens. On iOS it waits for the system prompt that names accounts.google.com and dismisses it. On Android it waits until Google's page covers the fixture's button, then taps Skip when the page has one and presses back when it does not, until the fixture is on screen again. On both it expects Google sign-in was cancelled. On Android it asserts on none of Google's text, because what Google Play services shows on an emulator with no Google account differs between images. The test proves that the hook reads the fixture's placeholder client IDs and reaches the native module, and that the module opens Google's sign-in and reports a cancel. It does not prove that a Google account can sign in, or anything about the ID token and the Clerk sign-in that follow. On iOS no Google page loads, because the cancel is at the system prompt.

The biometrics test has a settings file beside it, biometric-availability.settings.json, which turns biometric sign-in on for the instance, and run applies it before that file's test. The test expects biometric availability: biometric_authentication_unavailable on iOS, where the simulator has no Secure Enclave, and biometric availability: no_local_credential on Android, where the emulator has key storage and no stored credential. Both answers come from native code. On the standard settings the hook answers feature_disabled from JS, and the test fails. The test does not prove enrolling, storing, or signing in with a biometric, and neither answer changes when a biometric is enrolled on the device.

Four things keep a run steady on a slow device. host.tap taps again while a screen transition still covers the control, until the control is free or the tap times out. A launch opens the app a second time if the first try fails. Typing waits 40 ms between characters, and a test has 240 seconds.

No test is tagged known-bug. The close-button test of a full-screen AuthView in native-auth-view/opens runs with the rest, because #10079 is on main and in this branch. run leaves out a test with that tag unless --include known-bug is passed.

Each worktree gets one Clerk application in a workspace that holds nothing else, with the settings in src/core/instances/base.json. The CLI creates it with the Platform API key, which comes from the environment, a file, or a 1Password reference kept outside the repository. No file holds the application's secret key. Each command that needs it reads it from the Platform API and keeps it in memory until the command ends. After a run, the CLI searches the run directory for every secret the run used, and attach refuses to post a run that holds one.

The skill also configures e2e's built-in agent, for tests that act on the app or judge a screen with a model. It does so only when the machine has a Vercel AI Gateway key. The model is anthropic/claude-haiku-5.5, with openai/gpt-6-luna-fast as a backup that the gateway uses when the first model fails. No committed test uses the agent, and no workflow passes a key.

The skill installs e2e 0.18.0, @e2e-dev/mobile 0.10.0, @e2e-dev/github 0.4.0, and ai 7.0.128 with npm ci from its own lockfile, outside the pnpm workspace. The Verify Skill Tests job added to ci.yml runs the skill's 397 unit tests and tsc on Linux, with no device and no secret, when a pull request changes the skill or a path its tests read. On a Mac at this head, the 397 tests pass and tsc is clean.

Nothing in this pull request runs a test on a device in CI. The workflow in #10090 runs the 22 tests on a commit that contains this one.

The last run of that workflow that passed is run 37685636703, started by hand: 22 of 22 tests on iOS and 19 of 19 on Android, where three tests are for iOS only. No test needed its retry. The files it ran on included the borrowed-device code that is now in #10131, the older native-modules tests, and a fixture without the Google client IDs.

A later run on the present native-modules tests, run 37706627405, failed its Android job on the Google test. The CI emulator showed a Google Play services page that the back button does not cancel. The test now taps Skip on that page. No runner has run that step yet.

A local run on an iOS simulator passed 22 of 22 on an earlier head of #10090 that already had no borrowed-device code. Since that head, only the Google test and its feature file have changed, and the test's iOS steps are the same. Android was not run locally at that head.

The workflow from #10090 passed, started by hand on files identical to its present head, in run 37709561872: 22 of 22 tests on iOS and 19 of 19 on Android, where three tests are for iOS only. The Google test passed on both platforms on its first attempt. One iOS test, custom-flow-sign-in/complete, passed on its retry: the first attempt typed five of the six digits of the test code.

Not proven here:

  • A test with an agent step, since none is committed.
  • The Android half of the Google test on a CI runner, where Google's page has Skip.

Checklist

  • pnpm test runs as expected.
  • pnpm build runs as expected.
  • (If applicable) JSDoc comments have been added or updated for any package exports
  • (If applicable) Documentation has been updated

Type of change

  • 🐛 Bug fix
  • 🌟 New feature
  • 🔨 Breaking change
  • 📖 Refactoring / dependency upgrade / documentation
  • other: test tooling

🤖 Generated with Claude Code

@changeset-bot

changeset-bot Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 9c6673f

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
clerk-js-sandbox Ready Ready Preview Oct 8, 2026 12:53am UTC
swingset Ready Ready Preview Oct 8, 2026 12:53am UTC

Request Review

@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository YAML (base), Organization UI (inherited)
  • Review profile: ASSERTIVE
  • Plan: Team
  • Run ID: 7cd1585a-20df-484f-a882-1d122173b8d5
📥 Commits

Reviewing files that changed from the base of the PR and between e645244 and 7672257.

📒 Files selected for processing (1)
  • .claude/skills/verify-clerk-expo/test/github-report.test.ts
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 3 remain after this review.


📝 Walkthrough

Walkthrough

Adds the verify-clerk-expo skill for running and documenting @clerk/expo verification on iOS and Android. The implementation includes CLI commands, test-instance management, E2E execution and evidence handling, local simulator and emulator backends, and remote GitHub Actions sessions. It adds feature-specific specs and guides, shared skill access through a Cursor symlink, and CI jobs for skill checks and remote sessions.

Priority: ➖ Normal

Estimated code review effort: 5 (Critical) | ~120 minutes

Merge Risk: 🔵 Low · up to 76722

The verification tool can lose run evidence or encounter cleanup, duplicate-runner, and proxy-check failures in specific conditions. These are bounded workflow risks; merge is possible with owner awareness and follow-up.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 257 functions across 52 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the new Expo verification skill and its local-device fixture workflow. It omits borrowed-device support but remains directly related to the main changes.
Description check ✅ Passed The description directly explains the verification skill, CLI, fixture, device workflows, evidence capture, tests, CI scope, and limitations described by the changeset.
  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 160b44a to 729ecd1 Compare October 6, 2026 06:03
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 729ecd1 to 2d05137 Compare October 6, 2026 06:41
@mikepitre
mikepitre changed the base branch from mike/expo-verify-skill to mike/expo-verify-host October 6, 2026 16:30
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 2d05137 to 43ca8cc Compare October 6, 2026 16:55
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 43ca8cc to e31b477 Compare October 6, 2026 17:50
@mikepitre mikepitre changed the title test(expo): borrow a simulator or emulator on a CI runner for the verify skill test(expo): add a verify skill that drives the expo-native fixture on a local or borrowed device Oct 6, 2026
@mikepitre
mikepitre added this pull request to stack #10100 October 6, 2026 19:57
@mikepitre
mikepitre marked this pull request as ready for review October 6, 2026 19:59
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from caefb95 to 8647439 Compare October 7, 2026 17:01
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from b0d1508 to 249bffc Compare October 7, 2026 18:02
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 9b925c9 to e0f90bd Compare October 7, 2026 20:31
mikepitre and others added 7 commits October 7, 2026 20:50
…ring

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…roker, secrets, and evidence

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on per worktree

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

2 active deployments
Preview – swingset — 9c6673fa Deployed Oct 8, 2026 by vercel[bot]
Preview – clerk-js-sandbox — 9c6673fa Deployed Oct 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant