Repository navigation
feat: identify every gitops API request with a User-Agent - #78
Conversation
chris-garber-vapi
left a comment
There was a problem hiding this comment.
Aggressive review: 1 🟠, 3 🟡, 4 🟢, no blockers.
The main problem: the label doesn't reach child processes when a command isn't started through npm run, so PR check runs are counted as CI pulls (🟠 on user-agent.ts:61). The fix for that also keeps the label set bounded (🟡 on :63).
Every comment says what I checked. Every fix was tried on a copy of 674d8a1: tsc and npm test pass.
Not part of this diff, so no inline comment: under npm run promote, the promotion gate's simulation runs still send vapi-gitops-check/… (promotion-gate.ts:90). That means check-run counts include promotion gates.
1f982af to
908d125
Compare
674d8a1 to
5d90e98
Compare
|
@chris-garber-vapi on the gate label from your summary: done. Promotion gate runs now send 🤖 Generated with Claude Code |
chris-garber-vapi
left a comment
There was a problem hiding this comment.
Looks good, but it might be worth looking into if we should use X-Client-Source instead. I'm not sure how much value we lose from a debugging standpoint by overriding this field. Right now, the only values accepted by the api for X-Client-Source our dashboard and composer: I think it's worth extending that enum to include get apps, but also further extending it to parse x-client-version and x-client-action to capture some of the data that you've got here.
chris-garber-vapi
left a comment
There was a problem hiding this comment.
Reviewing a little more, I like these changes, I think we should do the X-Client-Source (and other headers) as a follow up.
Merge activity
|
## Value **V.A.L.U.E. tier:** project — PR 10 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)), the "check before deploy" step in promotion. > Stacked on #65 (the promotion partial-failure fix), now that #57–#64 have merged. This PR is the single gate commit. - **Problem:** promotion copies staging's reviewed files into production, but nothing checks that staging's agents still behave before they move on. Teams promoting dev → staging → prod need a behaviour gate between orgs, without new infrastructure. - **Who it affects:** multi-org gitops users (the promotion pipeline), who get a "check before deploy" step with one line of `promotion.yml`. Single-org users are unaffected. - **What changes:** - **`promotion.yml`** accepts `orgs.<slug>.check: <name>` (a slug), naming a `vapi-checks.yml` check. - **New `src/promotion-gate.ts`:** - **Validation before any transition:** the check must exist, and its `org` and `runOrg` must be the gated org, otherwise the run errors. `vapi-checks.yml` is required once any org is gated. - **Plan line:** what the gate would run, built offline. - **Live gate:** the check runs live and reduces to the worst target result. - **`src/promote-cmd.ts`:** in each transition, after the plan is built: - no changes skips the gate; - plan-only prints `check would run <name> in <org> (<n> simulations × <t> targets)`; - `--apply` runs the check (after the bindings refresh, before `promotionPlanApply` writes anything). Any non-pass throws `Promotion out of <org> blocked: check <name> <outcome> (<run url>)`. - A pass is cached per source org and dropped once a transition applies into that org. - `promotionCommandRun(args, overrides)` now takes `Partial<PromotionDeps>` (`childRun`, `checkRun`). - **`.github/workflows/promotion.yml`:** `timeout-minutes: 90` on the "Reconcile configured promotions" **step**, not the job, so the `if: always()` commit step (fixed in #65) still runs after a blocked or slow gate. - **Docs:** `promotion.example.yml` (a commented `check:`), a README "Check before promoting" section, and a pointer from "PR Checks". ## Evidence of value **The real gate, run live** in the owner's test org on the TEST-141 parity squad. - **Setup:** a scratch repo whose `promotion.yml` gates `parity` on check `core`, with pipeline `parity → parity-prod`. - **The run:** `promote --pipeline release --from parity --to parity-prod --apply`. - **The fake:** the child runner was faked, so bindings pulls were no-ops and the downstream `apply.ts` was recorded but not run. No second org was needed or touched. | Variant | Gate run | Result | Downstream apply | `resources/parity-prod/` | |---|---|---|---|---| | Degraded scheduler prompt | [7ed19587](https://dashboard.vapi.ai/simulations/run/7ed19587-d1ca-4d44-9232-7cdd15a50d67): 2 of 3 failed | `Promotion out of parity blocked: check core failed (https://dashboard.vapi.ai/simulations/run/7ed19587-…)` | **none** | **empty** (nothing written) | | Fixture as-is | [95470670](https://dashboard.vapi.ai/simulations/run/95470670-2861-45d3-a483-7a220fc3591a): 3 of 3 passed | promoted | `["parity-prod"]` | written; 20 applied paths recorded | The test org's resource counts were identical before and after both gate runs. **Tests:** `npm test` goes from 484 (#65) to 492 passing, and #68's golden promotion test passes unchanged. ## Testing plan - **`tests/promotion-gate.test.ts`** (6 tests, real git fixture, injected `childRun` / `checkRun`): - a pass applies; - failed and incomplete both block with the exact message, with no apply and the target untouched; - plan-only prints the line and runs nothing; - no changes skips the gate; - the three config errors (no `vapi-checks.yml`, unknown check, check in another org) stop before anything applies; - the pass cache: reused for two pipelines out of one org, and re-run after a transition applies into the gated org. - **No gate configured, no change:** with no `check:` in `promotion.yml` and an **invalid** `vapi-checks.yml` present, plan and `--apply` both succeed, `checkRun` is never called, and the plan output equals a pinned string. That string is exactly what #65's code (before the gate existed) prints for the same fixture, which I confirmed by running #65's `promote-cmd` on it. So the gate is invisible unless someone opts in. - **`tests/promotion.test.ts`:** `orgs.<slug>.check` is parsed, and a non-slug is rejected. - **Not tested:** - **A real two-org promotion:** only one test org was available. The downstream apply was faked, so the blocked case shows nothing written, and the pass case shows the apply was called. - **A GitHub Actions promotion run with a gate**, including the step timeout firing. - **Found while testing (pre-existing, out of scope):** promotion's dependency check rejects simulations that reference a **stock personality by UUID**, with "Referenced managed dependency is missing from source: personalities/a0000000-…". So a gated org's tests need local personality files until that's fixed. Stacked on #65. Refs TEST-141 ## After review The gate now refuses, when the config loads and before anything applies: - an unknown key under an org in `promotion.yml`, so a misspelled `check:` can't silently drop the gate; - `toolMocks: off` and `stripWebhooks: false`, because a gate runs in the real org, never a CI org; - a check `baseUrl` that differs from the org's `baseUrl` in `promotion.yml`, so the org's key only goes to the host promotion uses (the gate always uses that host); - a gate on an org that is last in every pipeline, where it would never run; - gated checks whose combined budget is over 300 minutes. It also fixes: - **Deadline:** each batch of 3 targets gets a full `timeoutMinutes`, so a check with more than 3 targets is no longer falsely blocked as incomplete. - **Step timeout:** raised from 90 to 330 minutes, as a safety net that no longer cuts short long ungated promotions. - **Tests:** the deadline and the worst-target rule are pure helpers with their own tests. The guide changes (blocks stop the whole run, simulation cost, the stock-personality limitation, accurate wording) are in #71. The `promote` User-Agent for gate runs is in #78. The block-report detail, fetch-stubbed gate test and deduplication are follow-ups. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Only `npm run sim` and `npm run check` identified themselves. Setup, pull, push, apply, promote, cleanup, rollback, call and audit sent Node's default `node` User-Agent, about a sixth of all api.vapi.ai traffic, so gitops usage beyond simulations couldn't be counted. - src/user-agent.ts: `vapi-gitops-<command>/<version>`, plus ` (ci)` when GITHUB_ACTIONS=true or CI is set (not false or 0). The command comes from a fixed list of this repo's commands, so a fork's own npm script names are never sent and npx is not a label; otherwise the entry script names it, else `cli`. The first process pins its label in VAPI_GITOPS_COMMAND, which spawned processes inherit, so the PR check's bindings pull is labelled check and `npm run apply`'s pull and push are labelled apply. sim and check keep their labels; promotion gate runs are labelled promote, apart from PR check runs. - Every fetch sends it, the GitHub status call included. tests/user-agent-coverage.test.ts checks each call on its own, scans src/ recursively, and checks api.ts against a local server. - cleanup-safety and new-file-gate tests sent about a dozen requests to the real api.vapi.ai per `npm test` (fake key, 401s), which the new User-Agent made visible in the request logs. tests/no-vapi-api.ts now loads before every test file: an unroutable base URL and no inherited real key, for the tests and every CLI they spawn. - how-it-works.md says exactly what the API sees; AGENTS.md says every request sends the header. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
5d90e98 to
6f0c59f
Compare

Value
V.A.L.U.E. tier: small — a behavior change: every API request now carries a new header. No blast-radius path.
npm run simandnpm run checkidentify themselves. Setup, pull, push, apply, promote, cleanup, rollback, call and audit send Node's default User-Agent,node, which was 126k of 728kapi.vapi.airequests in one hour this morning. So "how many deploys come from gitops, from CI or from laptops" has no answer.src/user-agent.tsbuildsvapi-gitops-<command>/<version>, plus(ci)whenCIorGITHUB_ACTIONSis set:<command>comes from a fixed list of this repo's commands, so a fork's own npm script names are never sent. The first process pins its label inVAPI_GITOPS_COMMAND, which child processes inherit:npm run applylabels its pull and pushapply, a promotion's applies arepromote, and the PR check's bindings pull ischeck. Promotion gate runs are labelledpromote, apart from PR check runs.simandcheckkeep their fixed labels, which existing simulation analytics already counts by.fetchto the Vapi API sends it:api.ts(push, pull, apply, promote), cleanup, setup, the interactive pickers, rollback, call, and push's direct fetch.cleanup-safetyandnew-file-gatesent about 12 requests pernpm testtoapi.vapi.ai, with a fake key, getting 401s. That breaks the repo's own rule that tests never call the real API, and with this PR it would have counted every fork's CI run as gitops usage. They now point at a dead local address.how-it-works.mdsays exactly what the API sees (and that there is no other telemetry).AGENTS.mdsays every request must send the header.Evidence of value
Live, through the Cloudflare request logs in Axiom (
cloudflare-logpush): a read-onlynpm run setup -- ua-check --resources noneagainst the test org, run from a scratch copy:vapi-gitops-setup/1.0.0vapi-gitops-test/1.0.0npm testruns before the fixBefore this PR, both rows would have been indistinguishable
nodetraffic.Tests:
tests/user-agent-coverage.test.tsreads everyfetch(insrc/and requires aUser-Agent. On the parent branch it lists 10 call sites without one; here it lists none. It also runsapi.tsagainst a local server withnpm_lifecycle_event=apply.tests/user-agent.test.tspins the format: fixed sim and check labels, npm script vs. entry script vs.cli, label cleaning, and the CI marker forGITHUB_ACTIONS=true,CI=trueandCI=1(but notfalse,0or empty).fetchtrap that records any request tovapi.aicaught 12 requests before the test fix and none after.Testing plan
npm test(527 tests) andnpx tsc --noEmitpass.1.0.0; bumping it is deliberately left out of this PR.callcommand's WebSocket audio connection isn't a Vapi REST request and doesn't carry the header.Refs TEST-141
After review
npxand fork script names fall through to the entry script, then tocli.fetchinsrc/does. The coverage test checks each call on its own and scans subfolders.tests/no-vapi-api.ts: loads before every test file, with an unroutable base URL and no inherited real key. With a pretend key exported and a fetch trap on, the suite makes zero requests to vapi.ai.how-it-works.mdmatches the code on(ci)and<command>.🤖 Generated with Claude Code