Repository navigation
perf(python): one-pass mirror of the parsed tree; map lookups in linkProject - #1853
Conversation
…hrough an index Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
index: three missing lookup indexes made large-repo queries and warm take minutes
…yable Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
build: export impact facts in the background once the graph is queryable
…k names the error Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
index now warms impact's facts in the background, so a query can arrive before edge.facts exists. context never exported them and died on a missing file. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…warm context: export the graph's facts before reading its edges
…at once The export had no lock: a query that found the stamp stale while another process was already writing the same facts (the background warm-up after index, a second query, several hooks at once) ran the whole export again beside it. Measured on an 8,619-file Java repository: the first query after index cost 193 s against 8.8 s warm, and every concurrent query paid the same again while thrashing the writer. The first comer now takes <facts>/.exporting and the others wait for its stamp, then answer from the facts it wrote; a lock whose writer is gone is taken over. tests/export_singleflight.py pins the wait, the takeover and the unlock. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…library staging A name no graph declares, reached only by text, whose lines include an import of it is the signature of an unstaged dependency — until now the answer was indistinguishable from an engine gap, and the fix is one flag away. One hint line under the [text] rows, only for undeclared names. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
impact/path: one facts export per graph, however many queries arrive at once
…s on 8,619 files Three wastes, found by timing each exported relation and profiling the poles standalone: - site_file() ran one has() and one paths lookup PER CALL, and the export calls it for every call edge: 638k round-trips over 319k edges. The paths map is now read once per graph (171 s -> 89 s). - via_base_rows probes field_access per single-target call site, and the table had no caller_id index: a 65,797-row scan 5,162 times. fa_caller lands with the other build-time indexes (89 s -> 63 s). - registrations() was computed twice in one export, once for the registration relation and again inside reg_key_fact (63 s -> 55 s). Every fact file is content-identical before and after (two differ in row order only, written from unsorted sets before this change too). Tests: java 329/329, python 287/287, facts_cache, latency, freshness, export_singleflight, front_door. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
impact/path: the facts export stops paying per-row SQL (171 s to 55 s on 8.6k files)
The Python stages each re-walked the tree-sitter tree through the JS<->C++ boundary, so a node's properties were marshalled once per stage; the parse stage of a large repository spent over half its time in those getters while the parse itself was negligible. materializePyTree now mirrors the tree in one cursor pass and every stage reads plain JS properties. Tree-sitter still parses every file. Two lookups in linkProject ran as linear scans inside loops (modules by hash per import record, bindings by hash per alias) and are now prebuilt maps with the same first-match semantics. Python analysis on two large corpus subjects drops 3.2x and 2.6x with byte-identical IR; the mirror is also A/B-asserted against the real tree node by node, which surfaced the one subtlety: an ERROR node absorbed during recovery is extra, so ERROR nodes read that flag from the real node. Parser version to 0.1.3 for the behaviour-neutral rebuild. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
The mirror moves to parsers/mirror-tree.ts, parameterized by each grammar's extras set; the Python module now only pins its own. Java and C# route their getRootNode through it (the real root remains available without a source string). JavaScript and TypeScript are untouched: their extractors run on the compiler AST, not tree-sitter. The port surfaced four tree-sitter subtleties the mirror now reproduces: named-sibling getters answer from anonymous nodes too; fieldNameForChild labels an extra on a field position while childForFieldName skips extras (field layout is re-read from the real node wherever an extra child makes the cursor's reporting untrustworthy); a file whose root carries an error reads every node's hasError from the real node, because a bare directive can report an error on itself with no visible ERROR child; and the C# mirror slices text from the BOM-stripped string parse() actually parsed. One deliberate behaviour change: Java's nested-annotation arguments were linked through a Map keyed by node OBJECT, and tree-sitter hands out a fresh wrapper per access, so the join hit only when a wrapper happened to be reused — 159 of 267 nested-annotation arguments linked on a large corpus subject. The mirror's stable identities link 260 of 267; the rest of the IR is byte-identical there, as it is on both Python subjects and the two C# subjects. Parse-stage wall time on large subjects: Java 66s -> 13s, C# 73s -> 18s. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
Extended to Java and C# (same branch): the mirror is now shared (
Validation: node-level A/B against real trees (django 400 files, jackson-scale subject all 1,386 files = 2.2M nodes, both C# subjects all 1,946 files — 0 consumer-visible mismatches), and all four language gate suites green: python 23/23, java 103 checks/0 failed, csharp 62/62 (torture manifest included), javascript 49/49. One deliberate behaviour change, Java only: nested-annotation argument linking used a Documented, consumer-unused divergences: |
… reads; solve non-main languages concurrently Three independent costs, one commit per stage touched: ts-relation-writer verified each finished relation by re-reading and re-decoding the whole file. The same rules — the header's field count and the consumer line-break alphabet — now run on each row's string as it is appended, in one allocation-free pass, and publish() proves the bytes arrived by comparing the byte count write() reported against the file's size, which is the defect the read-back existed to catch. The streaming verifier stays exported for the gate that exercises it. The root-program closure walk read and cheap-parsed every file to follow imports, then the extraction pass read every file again; the walk now hands its text over (consumed and released per file), and module resolution gets a ts.createModuleResolutionCache instead of re-probing node_modules per specifier. The all verb solved languages one after another although their solves share nothing. The largest language still solves alone first — its graph is the one a --progress caller publishes first — and the rest run in waves (AXIOMCODE_SOLVE_JOBS wide, default 2, 1 restores the strict line), each into its own intermediate, because run-souffle writes fixed names there. On a 1.1M-LOC TypeScript package, warm: stage 57.2s -> 50.9s with byte-identical IR; the removed double read was 12% of a cold parse. Serial and concurrent solves produce row-identical relations across all tables on a three-language tree. End to end on this repository (five languages), index wall time drops 37%. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
Third commit: the TS/JS-side and pipeline fixes. TypeScript parse (its extractors run on the compiler AST, so the mirror does not apply there — the costs were elsewhere):
Pipeline: End to end on this five-language repository: |
What
The Python extraction stages each re-walked the tree-sitter tree through the JS<->C++ binding, so every node's properties were re-marshalled once per stage. Profiling the parse stage on a 526k-LOC subject put 54.7% of samples inside the binding's property getters (
get typealone 16.1%);Parser.parseitself was 1%. Two lookups inlinkProjectadditionally ran as linear scans inside loops (modules by hash per import record, bindings by hash per alias).materializePyTreemirrors the freshly parsed tree into plain JS in one cursor pass; every stage then reads JS properties. Tree-sitter still parses every file — nothing about parsing changes.linkProjectprebuilds both maps with the same first-match semantics.Measurements (two large corpus subjects, warm, same machine)
Controls
diff -rqof the per-language output trees against the pristine base commit's parser).isExtra, so ERROR nodes read that flag from the real node (only reachable in files that failed to parse).npx tsx src/test/python-tests.ts: 23/23 gates pass.childForFieldName('alternative')onmatch_statementpierces into the match block; the stages iteratecase_clauses by type instead.