Repository navigation
v0.9.17: tables ledger, oracle integrations, mship improvements - #8853
Open
waleedlatif1 wants to merge 81 commits into
Open
waleedlatif1 wants to merge 81 commits into
waleedlatif1 wants to merge 81 commits into
Conversation
* feat(integrations): add saved Ramp and Vanta credentials * fix(integrations): harden credential renewal and connection setup * fix(integrations): make credential renewal waits abortable * docs(integrations): document token credential availability registration * improvement(integrations): refresh Ramp branding and simplify credential forms
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <bill@sim.ai>
…rface lint on deploy (#8803) * fix(workflows): lint unquoted string references in JSON fields and surface lint on deploy * fix(workflows): lint only the JSON fields the operation sends, and lint deploys as the acting user * fix(workflows): check JSON editors sent under a canonical param
* docs(library): update best-ai-agents-support-ticket-triage * Pi Babysit: address PR #8805 feedback --------- Co-authored-by: Sim Pi Agent <pi@sim.ai>
… 2026 (#8806) Co-authored-by: Sim Pi Agent <pi@sim.ai>
…ags (#8808) * chore(config): retire unused env vars and fully-rolled-out feature flags * chore(helm): bump chart version for removed TABLE_ROW_TTL value
* feat(oracledb): add Oracle Database integration * test(oracledb): remove retired block secret fields from masking inventory --------- Co-authored-by: Waleed Latif <walif6@gmail.com>
* fix(files): support arbitrary uploads and in-document links * fix(files): normalize previews and avoid oversized retries * fix(files): bound decoded text for encoded responses
…es a missing organization (#8809) * fix(billing): acknowledge Stripe webhooks whose subscription references a missing organization instead of retrying for days * test(billing): prove the not-yet-matching issuance branch by redelivery; drop the logging-only update catch * test(testing): a swapped price in the Stripe fake no longer reports the previous amount
* fix(db): retire unused keyword search projections * fix(db): retire keyword writers before schema push * fix(db): require approval before retiring push writers
* fix(google-ads): align reporting with current API access * fix(google-ads): preserve nullable inputs and redact retired credentials
* feat(dashboards): add empty-state graphic for the dashboard page Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA * improvement(dashboards): mark empty-state chart points as const Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA
…table definition row (#8765) * fix(tables): log row writes to the change log instead of locking the table definition row * fix(tables): lock the change log before the dev count reconcile, and tighten the held-row test * test(tables): name the renumbered change-log migration
* fix(tables): drop the per-table row-order lock from inserts Appends mint keys in a random slot after the last key, so concurrent appends never share a key and a batch stays contiguous. Positions are left best-effort and may repeat; the run dispatcher finishes a tied position before advancing its cursor. Positional inserts skip tied keys. Concurrent replaces stay serialized by the table's unique lock, which every replace already takes. * test(tables): order rows bytewise in the lockless-insert tests and stub the job queue locally
…ition source (#8815) * feat(analytics): attribute sign-ups and demo requests to their acquisition source - Record first and last marketing touch (campaign params, referring domain, landing path) in consent-gated first-party cookies on marketing and sign-in pages; attach them to user_created and the PostHog person - Capture $pageview on marketing routes, landing_demo_request_submitted, landing_demo_booked, external_sign_in_started, and email_type on identify - Add useCaptureWhenReady so view events captured on mount are no longer dropped before PostHog initializes - Include attribution in the demo-request sales notification - List the attribution cookies in the cookie policy * fix(analytics): cap attribution cookies at the consent grant, enforce stored field shapes, and lock sign-in buttons while pending
* feat(changelog): publish curated product updates * fix(changelog): harden feed rendering and media validation * improvement(changelog): present updates in a crawlable list * improvement(changelog): define the Sim editorial workflow Document scheduled discovery, durable candidate state, draft PR preparation, real-media review, and operational checks. Reuse the existing cached checkout for trusted Helm PR checks while preserving full history. * fix(changelog): preserve editorial decisions in the workflow plan * docs(changelog): refine product stories and define editorial quality * refactor(changelog): simplify presentation and validate model links
… leave-site prompt is answered (#8432) * fix(desktop-browser): ask the user before a page's alert, confirm, or leave-site prompt is answered * fix(desktop-browser): ask only about the user's own on-screen page, never the agent's * fix(desktop-browser): label frame dialogs with their own origin and cover reload leave prompts * fix(desktop-browser): preserve native dialog ownership and keyboard focus
…IPAA, GDPR, Data Residency, On-Prem, and VPC (#8822) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
* fix(mcp): add directory verification compatibility * test(mcp): verify ownership challenge over HTTP
* fix(chat): send entitlements and agent mode from the public chat API `/api/v2/chat` (what `sim chat` uses) stopped sending `entitlements` and `mode` in #8208, so CLI and API chats got no entitlement-gated tools and every CLI service refused them with "CLI services require agent mode". The route now computes entitlements per turn like the workspace chat and sends `mode: 'agent'`. Its test still mocked the old entitlements function and listed `mode` as a forbidden legacy field; both are updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): workflow tests as a workspace resource Workflow tests are a workspace resource whose source is a plain vitest file, `tests/<name>.test.js`, owned by the test (workspace_files.context = 'test'). Sim creates a test's metadata with the tests tool and writes its cases with the file tools; every write is collected in the sandbox and refused if the file does not load. - Runner: test files run in the isolated-vm sandbox against draft or deployed workflows. `runWorkflow` executes real runs; `mockBlock`, `mockTool` (Agent tool calls) and `spyOnBlock` reach blocks in the tested workflow and in every child workflow it runs, matched by name as each workflow starts. `.mockSampleOutput()` builds outputs shaped like the real block or tool. `toMatchRubric` asks a model judge for pass or fail. - Runs record live per-case progress, the source hash, and the deployment of every workflow they ran, so results show as out of date once the test or a workflow changes. - UI: Tests page and test page (Edit / Split / Preview over the file, the preview a dashboard of the selected run), a test resource type in chat, and a Tests sidebar entry behind the `workflow-tests` flag. - Owned files never open as file tabs in chat: only workspace files and chat uploads do. - Migration 0400 adds workflow_test and workflow_test_run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): pick a run from a dropdown and open what each run ran against The test page shows one run at a time, chosen from a run picker with status dots and Draft / Out of date chips. Case statuses use the Badge status chip. Each ran-against entry records one execution, so a draft row opens the workflow snapshot from that run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): type errors and tests broken by workflow tests - Narrow the test principal to the kinds workflow_tests.run admits before handing it to executeWorkflow. - Select progress with the latest-run rows, guard file upsert ids in the tab filter, and set the sandbox Event polyfills through Reflect. - Cover the tests tool in the management tool contract, expect content writes to reach test files, and stub test availability in the payload test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): address review on redaction, staleness, and the harness - Redact each run's resolved secrets from what returns to the sandbox (output, errors, mocked tool inputs); mocked tools get only declared params. - Custom blocks no longer receive the consumer's test hooks. - Draft runs go stale when the draft changes; children a run calls are recorded in ran-against. - Test cases commit in the same transaction as the source file write. - Harness: runWorkflow is rejected in suite hooks, a timed-out case stops the file, and expect.assertions/hasAssertions are supported. - Insert run rows in one statement and start each run's clock with its file; check bans before each workflow run; restrict owned-file access to Copilot delegation; validate names in the tool contract. - Delete soft-deletes the test file and removes the chat tab; a finished run shows its own cases; polling at 3s on a separate read bucket; list error state; store reset; tests stay in the org Add Resource picker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): reserve run slots, await pending assertions, keep dynamic tool args - Each test workflow run reserves and releases an execution slot. - A case waits for assertions it did not await and fails if one fails. - Mocked MCP and custom tools keep the arguments their schema declares. - A closed session refuses starts still awaiting their lookups. - Stable refresh callback; scroll fade on the results pane. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): keep test sources out of file tabs, refresh after runs, reopen tests - File-edit tool results mark a non-tab file `fileTab: false`, and the browser skips promoting it. - Idle test pages poll every 15s so runs started elsewhere appear; a Mothership run returns its tests as resource changes. - open_resource accepts test resources through an authorized read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): refresh test tabs after a run, keep saved edits successful - A finished Mothership run refreshes its tests instead of upserting tabs, so a test deleted mid-run does not come back. - A failed file-tab lookup after a saved edit opens no tab instead of reporting the edit as failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): starter source imports every test helper Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * build(tests): add @vitest/expect and @vitest/spy for the sandbox bundle The vitest-expect sandbox bundle builds from these packages; rebuilt with the Reflect-based event polyfills. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * build(tests): tell knip the sandbox bundle uses @vitest/expect and @vitest/spy Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): name MCP and custom tool mocks by server and title MCP tool ids embed the server's database id, which changes when a server is re-added or a workspace is forked, so a stale mock silently stopped matching and the real server was called. Tests now name workspace tools the way the workspace does: mockTool({ mcp: 'Server', tool: 'name' }) resolved per run (failing on an unknown or ambiguous server), and mockTool({ customTool: 'Title' }) matched case- and space-insensitively. Raw mcp- and custom_ ids are rejected; built-in catalog ids are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): reserve the run name, hold Run for unsaved edits, cancel judges on close - A test named "run" collided with the static run endpoint, so its detail page got a 405; the name is now reserved. - Run is disabled while the open editor holds edits the server has not saved (including a refused save), so a run never uses the previous source. - toMatchRubric model calls are aborted when the sandbox run ends, so a stopped test no longer keeps calling or billing the judge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): fail a run whose selected case names no test in the file A renamed or misspelled `only` path skipped every case, and the run was then saved as passing. The harness now rejects unknown names, so the run is recorded as an error with the names it could not find. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(providers): remove default agent tool-call iteration cap * docs(providers): note no execution timeout when billing is disabled
* fix(desktop): ask before accessing local files * fix(desktop): remember folder permissions across chats * fix(desktop): cancel pending file consent and recheck access * fix(desktop): enumerate approved directory descriptors * fix(desktop): reject replaced directory listings * fix(desktop): share pending folder consent decisions * fix(desktop): complete file consent and add full file access * fix(desktop): revalidate folder consent and preserve exact identities
* fix(pii): stop one half-emoji from failing log redaction A lone UTF-16 surrogate (e.g. an emoji cut by display truncation) made Presidio fail encoding its /redact_batch response, so mask-batch answered 500, each chunk was retried 8 times, and the default scrub mode replaced every string in the execution log with [REDACTION_FAILED]. - validate_pii: every string sent to Presidio goes out well-formed (toWellFormed, length-preserving U+FFFD); Presidio 4xx (other than 408/429) becomes PiiServiceRejectedError, unreachable / 502-504 becomes PiiServiceUnavailableError - mask-batch route: rejected -> 422 (never resent), unavailable -> 503 - mask client: optional failedChunkPlaceholder scrubs only the failed chunk and reports scrubbedCount; an outage (connection error or 502/503/504) stops sending queued chunks, a 500 does not - pii-redaction: scrub mode uses it and logs scrubbedCount, so one bad chunk no longer erases the whole payload; throw mode is unchanged - display-filters: truncate on a code-point boundary with an exact count - Presidio: any unencodable string in a request body is a 422, and 422 bodies no longer echo request input * test(pii): assert the outage breaker through masked output, not request counts * refactor(pii): drop unused error status, exports, and dead 422 escaping
…paging old ones (#8869) A billing Stripe sync that exhausted its retries stayed dead-lettered forever: requeue only sends it back through the same failure, retention never prunes a dead letter, and the hourly billing reconcile cron logged every one at ERROR on every run. - POST /api/v1/admin/outbox/[id]/resolve closes a dead-lettered billing Stripe event (dead_letter -> completed). last_error keeps who resolved it, why, and the last failure; the action is audited. Other domains count completed rows as delivered, so only billing Stripe events can be resolved. - The reconcile cron logs dead letters from the last two hourly runs at ERROR and older unresolved ones at WARN with the same identifiers; the dead-letter scan returns newest first so its cap never hides a new one. - The admin dashboard reports a resolved cancellation as resolved, not applied.
…s behind it (#8871) * fix(serializer): skip tool selection for trigger-mode blocks Trigger-mode blocks run through TriggerBlockHandler and never read a tool id, while their tool-mode sub-blocks (e.g. operation) are not serialized, so selecting a tool always threw and logged a fallback warning. * improvement(executor): memoize block output schemas per serialized block collectBlockData re-derived every block's output schema on each condition, function, agent, and reference evaluation, re-parsing an invalid agent response format and logging its full stack every time. Memoize the schema per immutable serialized block, attribute parse failures to the real block id, and log them at debug without the stack. * fix(webhooks): stop warning on accountless webhook credentials Service-account and managed-OAuth credentials (e.g. a custom Slack bot) have no OAuth account row, so the account-owner lookup always missed and warned on every webhook run. Skip the lookup for them and warn only when the credential or its account is genuinely missing. * improvement(webhooks): log Slack reactions.get missing_scope once per webhook A Slack bot without reactions:read fails reactions.get identically for every reaction event until it is reinstalled. Record that configuration state once per webhook per hour at info and keep the warning for other failures. * improvement(workflows): repair drifted subBlock types quietly Deployed snapshots are immutable and re-sanitized on every load, so a field whose declared type later changed (e.g. dropdown to combobox) warned on every materialization forever. Log repairs at debug when the stored type is one some registered block declares; keep warning for missing or unknown types. * improvement(catalog): skip custom block input derivation for tool reads Listing, reading, and executing a catalog tool resolved the catalog gate with every custom block's input fields, which loads and sanitizes each custom block's deployed workflow per request. Tool scope only needs the block types (every custom block exposes workflow_executor), so tool reads use the lightweight overlay rows; block reads still derive the inputs. * improvement(webhooks): debug-log missing idempotency ids for providers without one A generic webhook without an idempotency header or configured field has no delivery identifier by design, so warning on every delivery was noise. Warn only when the provider normally extracts an id from the body. * fix(flint): declare generate_pages items as an array param items was typed json, so the LLM schema advertised it as an object and dropped its item schema with a warning on every schema build. Declare it as an array of 1-10 page objects; block-provided JSON strings still parse through parseItems. Regenerate tool metadata and docs. * improvement(executor): skip JSON parsing for json inputs holding plain strings json-typed block inputs also accept plain strings (a file id, URL, or reference) that are kept as-is, so parsing them could only fail and warn. Parse only text that can be JSON; every valid JSON text still parses. * improvement(mothership): classify catalog routes and dedupe missing-provenance warnings Catalog responses (blocks, tools, connector types) carry no workspace data, so no producer records provenance for them and the per-request warning was noise. Skip them statically, and report a data-bearing route without a producer once per route per hour instead of on every request. Admission and recording behavior is unchanged. * fix(tokenization): resolve tiktoken encodings by model family encodingForModel only knows exact OpenAI names, so newer or provider-prefixed OpenAI ids fell back to cl100k_base (miscounting o200k models) and every unknown model built and cached its own copy of the rank table. Resolve exact names through js-tiktoken, newer OpenAI ids by family, everything else to cl100k_base, and cache one instance per encoding. * improvement(proxy): log blocked scanner requests at debug Every blocked request is an empty or tool user agent probing paths like /admin.php; the 403 is the whole response, so the per-request warning was pure ingest cost. * improvement(telemetry): downgrade best-effort collector failures to warn Forwarding to the external collector is best effort and the route still returns success, so a collector timeout or error is not an application error. * fix(admin): log mothership admin proxy failures and map upstream 5xx to 502 The proxy logged nothing, passed an upstream 5xx straight through as its own 500, threw on a non-JSON upstream body, and returned a silent 500 when the admin key was missing. Log the method, environment, endpoint, status, and a truncated upstream body; answer upstream 5xx and unparseable bodies with 502; and log the missing-key misconfiguration. * improvement(logs): move per-call success chatter to debug PII batch masking, function execution request/success, sandbox mount resolution, table row queries and limits, usage-limit statistics, DAG builds, and mothership tool routing logged an info line on every call. They are diagnostic detail, not events, so log them at debug. * chore(audits): shrink explicit-any baseline after tokenization cleanup * improvement(mothership): keep tool call routing decision at info It is the diagnostic for stale-catalog 'No handler for tool' dispatches and in-band double-dispatch races, so it stays visible in production. * fix(admin): keep upstream message and empty 2xx bodies in the mothership proxy A gateway 502 now carries the upstream's error or message field, which the admin UI shows, and an empty 2xx body (e.g. 204) passes through as an empty response instead of failing JSON serialization into a misleading 502. * improvement(tokenization): memoize model to encoding name resolution A non-OpenAI model threw and caught inside getEncodingNameForModel on every count. Cache the resolved encoding name per model id in a bounded LRU. * improvement(mothership): resolve catalog route patterns through the route table Derive each catalog pattern by matching its contract path against the generated v2 route table, so a renamed path parameter cannot drift from the pattern matchV2Route reports and silently restore the warning. * fix(webhooks): warn on missing idempotency ids unless the provider declares them optional Header-based providers (GitHub, GitLab, Shopify, Linear, Svix) have no body extractor, so keying the warning on the extractor silenced exactly the anomalous case of their delivery header going missing. Providers now declare deliveryIdOptional; only the generic webhook does, since its deliveries carry no identifier unless an idempotency field is configured. * refactor(logs): simplify log-noise fixes after review - Skip reactions.get entirely while a webhook's bot is known to lack the scope, instead of repeating a call known to fail. - Treat any intact sub-block whose stored type differs from the registry as drift, dropping the registry-wide type scan. - Move the JSON-candidate check next to isJSONString and trim once. - Make the catalog gate's cheap overlay rows the default; only block reads derive custom block inputs. - Collapse the admin mothership proxy's three copied handlers into one. * chore(logs): tighten comments and rename trigger-mode serializer test * test(mothership): guard sandbox catalog route classification against the route table Extract the catalog route set into isCatalogRoute, built lazily from the catalog contracts resolved through the generated v2 route table, and test it against real request paths so a route or parameter rename cannot silently restore the missing-provenance warning. * fix(webhooks): keep calling reactions.get after a missing-scope report Skipping the call for the cooldown window left a reinstalled Slack bot with empty reaction text for up to an hour. Keep the once-per-window log but always make the call. Block listing also uses the lightweight custom-block rows, since block summaries never read custom block inputs. * chore(logs): tighten log-noise changes after a line-by-line audit - Drop the blockId option threaded through getEffectiveBlockOutputs; it only labeled a debug log. - Name the log-dedupe windows and cache ceilings. - Build the sandbox catalog route set by converting contract paths to the route table's pattern syntax; the route-table test guards drift.
…8866) * improvement(changelog): open recordings in the shared lightbox * improvement(changelog): refresh product captures and released updates * improvement(changelog): write concise product announcements * docs(changelog): align the editorial body limit
… reconnect (#8868) * fix(oauth): record revoked refresh tokens durably and surface them as reconnect A refresh token the provider revoked (invalid_grant and kin) was only flagged in Redis for an hour, so every scheduled run kept failing with a generic error logged at ERROR three times, and the provider was asked again each hour, indefinitely. - account gains refresh_revoked_at/_code/_token_hash (expand-only, nullable). The record holds only while its hash fingerprints the stored refresh token, so a writer storing a new chain supersedes it; every reconnect path also clears it explicitly, since a provider can reauthorize without issuing a new refresh token. - The record is written only on rows still holding the rejected token, so a refresh that lost a race to a newer chain cannot mark the live credential revoked. Revocations no longer use the Redis dead flag, which now holds only app-registration faults. - A recorded revocation answers without the provider; one re-probe a day lets rejections that clear without a reconnect recover. A successful refresh clears it; a re-probe that fails for a passing reason keeps it. - refreshTokenIfNeeded throws CredentialRevokedError; token resolution returns code OAUTH_CREDENTIAL_REVOKED (POST and GET) and logs it once at WARN; the executor, tool and connector sites no longer log it at ERROR. * fix(oauth): guard revocation records against in-flight reconnects - Record a revocation only when the row's updated_at has not moved since the leader read it, so a reconnect that keeps the same refresh token while a rejected refresh is in flight is not undone. - Reread the row under the refresh lock, so a revocation another process just recorded is honored instead of calling the provider again. - A re-probe whose transient failure meets a moved chain reports no revocation. - Slack fan-out clears the record only for a chain the provider just issued, not when re-spreading a stored one. - A reconnect whose cleanup fails now fails the callback. - Revocation rejections log at WARN in oauth.ts too; the pure refresh error classifiers move to refresh-error-codes.ts so that client-reachable module can use them. - The connector path leaves the revocation log to executeSync. * fix(oauth): keep Slack connect fan-out unchanged and skip the cross-process check without Redis * test(oauth): clear only the fixture's Redis flags in the revocation suite * fix(oauth): clear only the recorded revocation on in-place relinks, with no Redis on the sign-in path
…rites (#8870) * fix(execution): tag deterministic admission rejections and throttle blocked-run logs Usage-limit, suspended-account, and missing-billing-account refusals now carry stable codes, so surfaces can tell a refusal that holds until a person acts from a transient one. Unattended surfaces can opt into throttleErrorLogs: each refusal of a workflow by the same gate records at most one execution error row per 15-minute window (Redis SET NX, in-process LRU without Redis, fail-open on Redis errors). A sender resending a refused delivery no longer writes an execution row, trace archive, and workspace_files row per attempt. A caller-supplied logging session is always completed. * fix(webhooks): acknowledge deterministic admission rejections for Telegram and Slack Telegram resends a non-2xx update until it succeeds or 24 hours pass, and Slack disables an app's event subscription once most deliveries fail, so answering a usage-limit refusal with 402 only loops. Providers opt in with acknowledgeAdmissionRejections and get an empty 200 with an ignored outcome; transient refusals (rate limit, concurrency, reservation outage) still fail so the sender retries. Generic webhooks keep the 402. Polling receives the raw refusal with its code. Webhook preprocessing throttles its error rows. * fix(webhooks): skip polls for over-limit payers and back off failing sources The poll orchestrator checks each workspace's payer once per tick and skips its webhooks while the payer is over its usage limit: nothing is fetched, marked seen, or counted as a failure, so items deliver once the payer is back under the limit and triggers are no longer auto-disabled over billing state. A webhook with consecutive poll failures waits 2^(n-1) minutes (capped at an hour) before the next fetch, and a source's Retry-After or FLOOD_WAIT_<n> is persisted and honored. RSS logs a source's 4xx once at warn, and an admission refusal mid-poll leaves the remaining items unseen. * fix(webhooks): stop Slack redelivering to deleted trigger paths and log them once A POST to a path with no webhook keeps its 404 but now carries x-slack-no-retry, sent unconditionally so it reveals nothing the 404 does not. The processor and route lines for an unknown path drop to debug, leaving the route handler's single client-error line. * fix(telegram): verify webhook deliveries with a per-webhook secret token Register a generated secret_token with setWebhook, store it as providerConfig.secretToken, and reject deliveries whose X-Telegram-Bot-Api-Secret-Token header does not match (401). Webhooks registered before this change have no stored secret and stay accepted until their next deploy registers one. A new registration reuses the active deployment's secret for the same bot so deliveries keep verifying during cutover. Drops the stale empty User-Agent warning: the proxy exempts webhook trigger routes from the empty-UA block. * refactor(webhooks): declare the Telegram admission opt-in on its handler * fix(execution): only treat billing and account refusals as deterministic An unreadable usage ledger fails closed as exceeded; it said nothing about the payer, yet it was tagged USAGE_LIMIT_EXCEEDED and acknowledged-and-dropped for Telegram and Slack. It now stays untagged and retryable. Reservation headroom denials clear as in-flight runs settle, so they leave the deterministic set too. The blocked-run log claim drops its in-process fallback: usage refusals only happen on hosted billing deployments, which run Redis, and without Redis every refusal records its row. A Redis failure logs at debug. * fix(webhooks): keep a fan-out target's retryable failure visible past a dropped refusal An acknowledged admission refusal counted as an acknowledgment, so a Slack or path fan-out answered 200 even when another target failed and needed the sender to retry. Dropped targets (block missing, acknowledged refusal) now answer 200 only when no other target failed. * fix(telegram): match the active bot through env-var token references Active rows store the bot token as authored, often a {{VAR}} reference, while subscription calls receive it resolved, so the comparison never matched: every deploy minted a fresh secret (a candidate that never activated left the bot rejected by the active row), and retiring an old version could delete the webhook the active version still used. Stored tokens are now resolved against the background webhook env before comparing. * fix(webhooks): skip polls only after a recorded refusal and back off source failures The per-tick payer pre-check read billing attribution and the usage ledger for every polled workspace, including healthy idle ones. A deterministic admission refusal of a polled event now records the workspace in Redis for five minutes, and the orchestrator skips only those workspaces in one MGET; healthy payers cost no billing reads, and billing-disabled deployments skip the mechanism. Backoff now follows source fetch failures only, tracked in providerConfig and stamped from the failed poll's start, so transient item-processing refusals never back a webhook off. Poll state keys are system-managed so deploy change detection ignores them. * refactor(webhooks): route every poller's source failures through one backoff Every poller's outer catch now records a source failure, so a failing Gmail, Outlook, IMAP, Drive, Sheets, Calendar, or HubSpot source backs off like RSS instead of only RSS. markWebhookSuccess clears the backoff in its existing reset write, the window uses the shared jittered backoff, and the orchestrator asks a boolean isPollBackedOff. Smaller cleanups: one isDroppedDispatch predicate for the Slack and path fan-outs, explicit precedence for a polled refusal's code, a typed RSS refusal error instead of a flag, Telegram resolves only the stored bot token, and PollOutcome lives with the polling types. * chore(webhooks): tighten poll comments and backoff tests Poll outcome docs describe what a skipped poll actually does, source failures keep logging the full error object as the pollers did before, the backoff table pins the clock past the poll start so it proves the window is anchored there, and the RSS rate-limit test drives a Retry-After header. * fix(webhooks): stop every poller's batch on a deterministic admission refusal Only RSS stopped at a refused item; Gmail, Outlook, and IMAP advanced their cursors past refused emails, and every poller counted the refusals toward auto-disable. A shared PollAdmissionRefusedError now leaves the idempotency callback, stops the batch, and returns skipped before any cursor update or failure count; items that already ran replay as idempotent no-ops. Source backoff goes back to RSS only, where the rate-limited feed was: the other pollers' fetch helpers do not carry status or Retry-After, so routing their failures through it would back off on a guess. The block-missing 404 also tells Slack not to redeliver. * fix(webhooks): never replay completed poll work or mask a retryable fan-out failure A poller stops its batch on a deterministic refusal only while nothing in the batch has completed; once an item has run, the refusal is an ordinary item failure, so the poller saves its completed work exactly as before and no completed event can replay after the idempotency window. In a multi-target delivery a missing block's no-retry 404 no longer stands in for a target that failed and needs the sender to retry. A source's Retry-After is counted from its answer rather than the poll's start. The RSS backoff keys are cleared by RSS's own state write, so other pollers' success path is unchanged, and two fields nothing reads are dropped. * fix(execution): drop the unreachable billing-account admission code A workspace without a billing account fails inside payer resolution and takes the retryable attribution-error path; the branch that tagged BILLING_ACCOUNT_REQUIRED only ran for an attribution with no actor, which system attribution never produces. The branch goes back to its staging form and the deterministic set keeps the usage limit and suspended accounts. * chore(webhooks): key blocked-run claims by gate and centralize the polling utils mock Two gates that fail without a code (a ban lookup error and a usage lookup error) no longer share one throttle claim, so neither hides the other's row. The polling utils module gets one central mock in @sim/testing, replacing the partial importOriginal mocks and the hand-rolled factory in the table trigger test. * fix(telegram): match the active bot with the env the caller resolved its token with Subscription creation resolves the incoming bot token with the deployer's env, while cleanup resolves with the background env; the active-row matcher now uses the same env as its caller, so a bot referenced through a personal variable still reuses the active secret and is not deleted from under the active deployment. Tests: the idempotency service gets one central mock in @sim/testing, used by every test that mocked it locally, and the new tests import single factory and mock files instead of the @sim/testing barrel. * chore(testing): stub IdempotencyService.createWebhookIdempotencyKey in the central mock
…ons they declare (#8877) * improvement(sim-cli): commands reach the API only through the operations they declare Hand-written commands called SimClient by path and declared nothing, so the command inventory could not mark one Mothership-unavailable when it calls a refused route. Every command now gets its client from callsOperations: calls name an operation, route and method come from the operation table, and the client is typed to the declared operations. * test(sim-cli): assert declared-operation routing at the transport * test(sim-cli): assert directory requests at the transport
…dary (#8872) * improvement(logs): log each execution failure once at its owning boundary A single external tool failure was logged at ERROR up to eight times as it propagated from the tool layer through the block executor, the engine, execution-core, and the trigger surfaces. Each boundary now calls logFailureOnce, which skips a failure an inner boundary already logged and sets severity by attribution: the author's input, configuration, or code at info, a third-party 4xx or 5xx at warn, and Sim's own faults (database, retryable setup, hosted-key rejection, anything unattributed) at error. The logged mark and attribution live in WeakMap/WeakSet side tables keyed by the thrown value, walked through the cause chain, and carried on the flattened tool failure's output so they survive a handler rebuilding a failed tool result as a new error. Run logs and trace spans are unchanged. * test(logs): pin the logged mark and attribution at the block executor Asserts the block executor's thrown error is already logged and attributed (database failure internal, user Stop user), switches the tool-boundary test to the central jsonResponse helper, and updates exact-argument log assertions for the added failureKind and executionId fields. * fix(logs): scope the logged mark to one propagation and attribute in-process failures Never mark a raw thrown value as logged: a persistent fault that rethrows one object (a rejected dynamic import, a memoized rejected promise) was logged once per process and then silenced everywhere. Only carriers created during the propagation (the flattened tool failure output, the block error, a handler's rebuilt error) carry the mark, and execution-level boundaries scope a raw value's mark to their execution id. Handlers that rebuild a failed tool result now call adoptToolFailure instead of relying on the result output being copied by reference. The child workflow and agent handlers no longer log failures the block executor logs with the block's and run's identity. In-process operations are attributed to Sim (4xx user, 5xx internal) rather than to a third party, only external HTTP failures log their response body, a missing required field is the only serializer refusal at info, and the size-limit and JSON-parse paths no longer log twice. * refactor(logs): build failure log metadata lazily and attribute author failures by type logFailureOnce takes its metadata as a thunk so a skipped boundary does no secret projection, adds the execution id itself, and returns the attribution it logged with. UserFailure attributes author-caused failures by type (missing required fields, boundary-safe custom block refusals, no start block) wherever they surface, replacing a try/catch at one serializer caller. One inheritFailureMarks carries both marks across a boundary that drops cause, the hosted-key check reuses classifyHostedKeyFailure, the Pi backends share one toolResultError, and the sandbox and Function route no longer log author code failures the tool boundary already logs. * test(logs): attribute the serializer refusal at its owner and pin block log redaction Moves the missing-required-fields attribution check to the serializer that throws it, drops an execution-core case whose internal row passed by default, and covers the block failure line's secret projection now that it, not the Agent handler, logs provider errors. Tightens comments the diff added. * improvement(logs): drop casts and test-only exports from the failure log Reads status and cause without casts, keeps the thrown missing-required-fields error's name and stack unchanged, ignores an empty execution id when scoping a logged mark, unexports wasFailureLogged and FailureKind (tests observe the outer boundary through logFailureOnce instead), and shrinks the explicit-any baseline for the Agent handler. * fix(logs): close the remaining double logs found in review Carries failure marks through the scrubbed Pi error and normalizes a non-Error node failure before logging so the engine sees its mark; drops the condition handler's and response-size handler's own error lines in favor of the owning boundary; logs MCP failures once and marks their output; passes the block id to the API and Function tool calls so the tool line carries it; attributes custom tool parameter validation and condition expression errors to the author; checks the whole cause chain for a retryable setup failure before honoring a mark; and restores the sandbox failure logs, whose adapters cannot yet tell a provider exception from a non-zero exit. * fix(logs): keep run identity on every remaining failure line Fixes the type error on the external failure body, attributes a provider error payload on a 2xx to the provider, and adds workflow, execution, and block ids to the MCP failure lines that replace the block executor's. The condition batch carries its failed result's marks and passes its block id. Pi GitHub calls carry no run identity, so their errors now carry only the tool layer's attribution and the block executor still logs them with ids. Restores the HITL notification warning, which covers soft failures the tool layer never logs, and attributes refused tool and proxy URLs to the author. * fix(logs): attribute a revoked OAuth credential to its owner A CredentialRevokedError anywhere in the cause chain classifies as a user failure, so the boundaries that log it after token resolution's WARN write INFO rather than ERROR. * fix(logs): keep one logged mark per execution and attribute unparseable responses A raw value logged at an execution boundary now remembers every execution that logged it (bounded), so overlapping runs sharing one persistent rejection each log it once. A 2xx body that is not JSON is the endpoint's failure, not Sim's.
* improvement(app): show branded wordmark while loading * fix(app): preload workspace data alongside branding
…data-typed next/script (#8875) * fix(docs): render the FAQ JSON-LD as a native script tag, and forbid data-typed next/script The FAQ component rendered its FAQPage JSON-LD through next/script with the default afterInteractive strategy, which returns null on the server and injects the tag from a client effect, so the served HTML never carried it. Render a native <script> the way #6763 fixed the other docs JSON-LD blocks; serializeJsonLd already escapes '<'. check:source-text now also fails on a next/script element in apps/** whose type is not JavaScript, naming the native-script fix. * fix(audits): match only the real type prop, fail closed on expression types, and scan MDX pages * fix(audits): recognize default-as next/script imports and skip MDX code examples * fix(audits): keep JSX template literals intact when skipping MDX inline code
…ode; fix the PPTX blend-group call they hid (#8876) * chore(audits): drop dead ESLint directives and the redundant check:dead-code script - Remove all 89 `eslint-disable` comments. Nothing runs ESLint (no eslint dependency or config in any workspace; Next 16 `next build` has no lint step), so they suppressed nothing. Reasons they carried are kept as plain `//` whys. - `check:comment-hygiene` now rejects `eslint-disable`/`eslint-enable` comments so they cannot return. - Remove `check:dead-code`: `check:unused-exports` already runs knip with the same config and a superset of its issue types, and `check:audits` skipped the alias. - Ignore unused exports in the generated `apps/docs/components/icons.tsx` (a verbatim copy of the sim icon set, whose export surface is ratcheted at the source) and shrink the unused-exports baseline by its 66 entries. * fix(pptx): stop calling the appended rect blend group as a function * docs(skills): name both ESLint directives the comment audit rejects * docs: name both ESLint directives in the comment rule * test(pptx): cover the rect gradient blend-group path
…adline (#8878) * fix(tools): stop provider timeout params from becoming the request deadline The request transport and the internal-operation path read params.timeout as a millisecond deadline for every tool. Twilio make_call, New Relic NRQL, Apify (3 tools), Daytona (2 tools), and Trigger.dev waitpoint tokens declare their own timeout param in seconds or as a duration, so a 60-second setting aborted the call after 60 ms. A declared timeout param is now the deadline only when the tool sets timeoutParamIsDeadline (http_request, firecrawl_map, firecrawl_parse); callers can still bound tools that declare none. No param ids change. Redis and Upstash coerced params with Number() inside tools.config.tool, which runs at serialization on the serialized params object, turning <Block.output> references into NaN. The coercions now run in tools.config.params. Guardrails: check-block-registry rejects coerced assignments to params inside an inline tools.config.tool; check-tool-param-reachability rejects a method param on a fixed-verb external tool, which the transport would send as the HTTP verb. * fix(audits): catch coercions under fallbacks in tools.config.tool, and clarify declared timeout params * fix(audits): treat every value-deriving expression as a selector coercion * fix(audits): reject compound assignments and increments on params in selectors
…through wrappers (#8885) - v1 admin: unauthorized/forbidden/notFound/badRequest/internalErrorResponse -> admin*Response; errorResponse module-local; conflict/notConfiguredResponse also admin-prefixed - files createErrorResponse -> createFileErrorResponse; workflows createErrorResponse -> createCodedErrorResponse (+ central mock) - import VFS segment codecs from @/lib/vfs/path; delete the mothership pass-throughs and the dead canonicalizeVfsPath - search-replace resource group key uses sortObjectKeysDeep instead of a localeCompare serializer
… discovers them (#8882) * chore(audits): name generated-artifact checks check:<x> so run-audits discovers them Rename tool-metadata, deployment-config, integration-catalog, docs, agent-stream-docs and docs-manifest checks to check:<x> (generators to generate:<x>, matching the existing generate:openapi/check:openapi pairs). run-audits now derives every audit from the check:* prefix, so the hand-kept EXTRA_AUDITS list is gone, and check:docs-manifest (local fs only, ~1s) joins check:audits instead of its own CI step and CLAUDE.md gate line. The <x>:check suffix stays only for checks this job cannot run (sibling copilot repo, Helm, per-machine covers). Add module headers to check-monorepo-boundaries, check-realtime-prune-graph and check-openapi saying what each protects and how to fix a finding. * docs(ship): regenerate tool metadata in Phase A; drop the actionlint comment with no command
…8884) * chore(naming): rename 41 baselined files to the file-name convention Kebab-case the guardrails validators and azure-blob destination, drop folder stutter in lib/logs, hooks/queries/oauth, emcn charts/chip/popover/tooltip/ tab-strip/calendar and workflow-renderer note/subflow/workflow-block, and rename _polyfills/_nodemailer. Updates every importer, vi.mock string, package exports/imports/sideEffects target and doc reference; fixes the anys and redundant non-null assertions in renamed files. check:file-names baseline 171 -> 130. * docs(workflow-renderer): point the handle-position comment at the renamed views
* docs(agents): align agent guidance with code and enforced checks - contracts: export the contract; export a schema or type only when imported - testing: mock-assertion ban matches test-audit; .dom.test convention; runner commands - imports: biome organizeImports owns order - stores: devtools for new stores; reset() + registerUserDataReset - add apps/sim/CLAUDE.md and (landing)/AGENTS.md symlinks; name db-migrate skill - skills: safeUrlPathSegment for path segments, hmacSha256Hex and bounded reads in trigger templates, hosted-key per_request/enabledWhen, stale identifiers fixed, pressure markers removed - CONTRIBUTING: commit types and format match history * docs(add-model): keep the check:/generate: script names from staging * docs(emcn-design-review): list only the Chip variants its props accept
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Companion: simstudioai/mothership#656 (merge this first; then re-run sim-sync on #656)