Skip to content

perf(python): one-pass mirror of the parsed tree; map lookups in linkProject - #1853

Merged
swapnilpaliwal-sd merged 16 commits into
apps/integration-0.1.9from
perf/parser-one-walk
Oct 6, 2026
Merged

swapnilpaliwal-sd merged 16 commits into
apps/integration-0.1.9from
perf/parser-one-walk

Conversation

@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor

What

The Python extraction stages each re-walked the tree-sitter tree through the JS<->C++ binding, so every node's properties were re-marshalled once per stage. Profiling the parse stage on a 526k-LOC subject put 54.7% of samples inside the binding's property getters (get type alone 16.1%); Parser.parse itself was 1%. Two lookups in linkProject additionally ran as linear scans inside loops (modules by hash per import record, bindings by hash per alias).

  • materializePyTree mirrors the freshly parsed tree into plain JS in one cursor pass; every stage then reads JS properties. Tree-sitter still parses every file — nothing about parsing changes.
  • linkProject prebuilds both maps with the same first-match semantics.
  • Parser version 0.1.2 -> 0.1.3.

Measurements (two large corpus subjects, warm, same machine)

subject python stage before after speedup
subject A (526k LOC) 211.7s 66.9s 3.2x
subject B (638k LOC) 160.9s 62.9s 2.6x

Controls

  • IR is byte-identical on both subjects (full diff -rq of the per-language output trees against the pristine base commit's parser).
  • Node-level A/B harness asserts every mirrored property (type, spans, counts, text, flags, fields) against the real tree: 2.3M nodes, 0 mismatches. It caught the one real subtlety: an ERROR node absorbed during recovery is isExtra, so ERROR nodes read that flag from the real node (only reachable in files that failed to parse).
  • npx tsx src/test/python-tests.ts: 23/23 gates pass.
  • Known accepted divergence (unused by any stage): tree-sitter's childForFieldName('alternative') on match_statement pierces into the match block; the stages iterate case_clauses by type instead.

swapnilpaliwal-sd and others added 14 commits October 1, 2026 13:46
…hrough an index

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
index: three missing lookup indexes made large-repo queries and warm take minutes
…yable

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
build: export impact facts in the background once the graph is queryable
…k names the error

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
index now warms impact's facts in the background, so a query can arrive before
edge.facts exists. context never exported them and died on a missing file.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…warm

context: export the graph's facts before reading its edges
…at once

The export had no lock: a query that found the stamp stale while another
process was already writing the same facts (the background warm-up after
index, a second query, several hooks at once) ran the whole export again
beside it. Measured on an 8,619-file Java repository: the first query
after index cost 193 s against 8.8 s warm, and every concurrent query
paid the same again while thrashing the writer.

The first comer now takes <facts>/.exporting and the others wait for its
stamp, then answer from the facts it wrote; a lock whose writer is gone
is taken over. tests/export_singleflight.py pins the wait, the takeover
and the unlock.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…library staging

A name no graph declares, reached only by text, whose lines include an
import of it is the signature of an unstaged dependency — until now the
answer was indistinguishable from an engine gap, and the fix is one flag
away. One hint line under the [text] rows, only for undeclared names.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
impact/path: one facts export per graph, however many queries arrive at once
…s on 8,619 files

Three wastes, found by timing each exported relation and profiling the
poles standalone:
- site_file() ran one has() and one paths lookup PER CALL, and the
  export calls it for every call edge: 638k round-trips over 319k
  edges. The paths map is now read once per graph (171 s -> 89 s).
- via_base_rows probes field_access per single-target call site, and
  the table had no caller_id index: a 65,797-row scan 5,162 times.
  fa_caller lands with the other build-time indexes (89 s -> 63 s).
- registrations() was computed twice in one export, once for the
  registration relation and again inside reg_key_fact (63 s -> 55 s).

Every fact file is content-identical before and after (two differ in
row order only, written from unsorted sets before this change too).
Tests: java 329/329, python 287/287, facts_cache, latency, freshness,
export_singleflight, front_door.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
impact/path: the facts export stops paying per-row SQL (171 s to 55 s on 8.6k files)
The Python stages each re-walked the tree-sitter tree through the JS<->C++
boundary, so a node's properties were marshalled once per stage; the parse
stage of a large repository spent over half its time in those getters while
the parse itself was negligible. materializePyTree now mirrors the tree in
one cursor pass and every stage reads plain JS properties. Tree-sitter still
parses every file.

Two lookups in linkProject ran as linear scans inside loops (modules by hash
per import record, bindings by hash per alias) and are now prebuilt maps
with the same first-match semantics.

Python analysis on two large corpus subjects drops 3.2x and 2.6x with
byte-identical IR; the mirror is also A/B-asserted against the real tree
node by node, which surfaced the one subtlety: an ERROR node absorbed
during recovery is extra, so ERROR nodes read that flag from the real node.

Parser version to 0.1.3 for the behaviour-neutral rebuild.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
The mirror moves to parsers/mirror-tree.ts, parameterized by each grammar's
extras set; the Python module now only pins its own. Java and C# route their
getRootNode through it (the real root remains available without a source
string). JavaScript and TypeScript are untouched: their extractors run on
the compiler AST, not tree-sitter.

The port surfaced four tree-sitter subtleties the mirror now reproduces:
named-sibling getters answer from anonymous nodes too; fieldNameForChild
labels an extra on a field position while childForFieldName skips extras
(field layout is re-read from the real node wherever an extra child makes
the cursor's reporting untrustworthy); a file whose root carries an error
reads every node's hasError from the real node, because a bare directive
can report an error on itself with no visible ERROR child; and the C#
mirror slices text from the BOM-stripped string parse() actually parsed.

One deliberate behaviour change: Java's nested-annotation arguments were
linked through a Map keyed by node OBJECT, and tree-sitter hands out a
fresh wrapper per access, so the join hit only when a wrapper happened to
be reused — 159 of 267 nested-annotation arguments linked on a large
corpus subject. The mirror's stable identities link 260 of 267; the rest
of the IR is byte-identical there, as it is on both Python subjects and
the two C# subjects.

Parse-stage wall time on large subjects: Java 66s -> 13s, C# 73s -> 18s.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Extended to Java and C# (same branch): the mirror is now shared (parser/src/parsers/mirror-tree.ts), parameterized by each grammar's extras; JS/TS are untouched (their extractors run on the compiler AST, not tree-sitter).

subject stage before after speedup IR vs pristine
Python, 526k LOC 211.7s 66.9s 3.2x byte-identical
Python, 638k LOC 160.9s 62.9s 2.6x byte-identical
Java, 318k LOC 66.2s 13.3s 5.0x identical except the one column below
C#, two subjects 73.2s 17.6s 4.2x byte-identical

Validation: node-level A/B against real trees (django 400 files, jackson-scale subject all 1,386 files = 2.2M nodes, both C# subjects all 1,946 files — 0 consumer-visible mismatches), and all four language gate suites green: python 23/23, java 103 checks/0 failed, csharp 62/62 (torture manifest included), javascript 49/49.

One deliberate behaviour change, Java only: nested-annotation argument linking used a Map keyed by node OBJECT, and tree-sitter hands out a fresh wrapper per access, so the join hit only when a wrapper instance happened to be reused — 159/267 nested-annotation args linked on the large Java subject. With the mirror's stable identities it links 260/267. Every other Java relation is byte-identical.

Documented, consumer-unused divergences: childForFieldName('alternative') piercing on Python match_statement, and named-sibling resolution around ERROR/hidden boundaries in C# (no C# stage reads named siblings).

… reads; solve non-main languages concurrently

Three independent costs, one commit per stage touched:

ts-relation-writer verified each finished relation by re-reading and
re-decoding the whole file. The same rules — the header's field count and
the consumer line-break alphabet — now run on each row's string as it is
appended, in one allocation-free pass, and publish() proves the bytes
arrived by comparing the byte count write() reported against the file's
size, which is the defect the read-back existed to catch. The streaming
verifier stays exported for the gate that exercises it.

The root-program closure walk read and cheap-parsed every file to follow
imports, then the extraction pass read every file again; the walk now
hands its text over (consumed and released per file), and module
resolution gets a ts.createModuleResolutionCache instead of re-probing
node_modules per specifier.

The all verb solved languages one after another although their solves
share nothing. The largest language still solves alone first — its graph
is the one a --progress caller publishes first — and the rest run in
waves (AXIOMCODE_SOLVE_JOBS wide, default 2, 1 restores the strict line),
each into its own intermediate, because run-souffle writes fixed names
there.

On a 1.1M-LOC TypeScript package, warm: stage 57.2s -> 50.9s with
byte-identical IR; the removed double read was 12% of a cold parse.
Serial and concurrent solves produce row-identical relations across all
tables on a three-language tree. End to end on this repository (five
languages), index wall time drops 37%.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Third commit: the TS/JS-side and pipeline fixes.

TypeScript parse (its extractors run on the compiler AST, so the mirror does not apply there — the costs were elsewhere):

  • relation CSVs were verified by re-reading and re-decoding the finished file; the same width + line-break rules now run per row at append time, allocation-free, and publish() checks written-bytes against file size (the torn-write defect the read-back existed for). The streaming verifier stays exported for its gate.
  • the root-program closure walk read + cheap-parsed every file to follow imports, then extraction read everything again; the text is now handed over (consumed per file), and resolveModuleName gets a ModuleResolutionCache instead of re-probing node_modules per specifier.
  • warm A/B on the 1.1M-LOC TS package: stage 57.2s → 50.9s, IR byte-identical; the removed double read was ~12% of a cold parse. TypeScript suite 49/49.

Pipeline: all solved languages strictly in sequence. The largest language still solves alone first (publish-first behaviour unchanged), the rest now run in waves (AXIOMCODE_SOLVE_JOBS, default 2; 1 restores the old line), each with its own intermediate because run-souffle writes fixed file names there. Serial vs concurrent produce row-identical relations across all 255 tables on a three-language tree.

End to end on this five-language repository: axiomcode index wall time 207.5s → 129.9s (−37%).

@swapnilpaliwal-sd
swapnilpaliwal-sd merged commit d515f5d into apps/integration-0.1.9 Oct 6, 2026
12 checks passed
@swapnilpaliwal-sd
swapnilpaliwal-sd deleted the perf/parser-one-walk branch October 6, 2026 02:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant