Skip to content

Clip selected rows before nested VARIANT assembly - #179

Merged
platypii merged 1 commit into
masterfrom
codex/bounded-variant-decoding
Sep 17, 2026
Merged

platypii merged 1 commit into
masterfrom
codex/bounded-variant-decoding

Conversation

@philcunliffe

Copy link
Copy Markdown
Contributor

Small row-range reads currently assemble nested VARIANT values for the entire covering physical range before slicing the result. Compact repeated tool definitions can therefore expand into hundreds of MiB even when the caller requests only a few rows.

Intersect the requested selection with the common child-column range before nested assembly. Preserve absolute skipped-row coordinates and alignment when physical children start at different offsets. The regression places malformed VARIANT values outside the selection and proves they never reach the decoder.

Validation: 665 tests passed; lint, declaration build, and git diff --check passed. With coordinated HypAware consumer windowing, a previously failing export completed under a 128 MiB V8 heap cap. A 4x row-count stress run kept sampled live heap near 36 MiB.

CPU and memory: nested expansion is proportional to the selected physical range, avoiding repeated full-group object/string allocation. Physical page decoding, scalar dictionaries, encoded buffers, and individually large values retain their existing memory costs.

@platypii platypii left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correctness looks good: clipping preserves child-column alignment and row offsets. Additional local tests exhaustively exercised selection boundaries and independently aligned child pages over a six-row group, including empty selections, and checked both object and array output. Selecting 10 of 10,000 VARIANT rows decoded exactly 10 values and matched the corresponding full-read results. All 667 tests, lint, and type checks passed with these additional tests in the working tree; those tests are not yet included in this PR.

Performance: a local Node v26.5.1 microbenchmark using 10,000 nested VARIANT rows (79-byte fixture values), warmup, and 50 alternating runs measured median assembly plus row-conversion time against the previous clipping behavior:

  • Select 10 rows: 6.405 ms → 0.012 ms, about 550× faster.
  • Select 1,000 rows: 7.127 ms → 0.691 ms, about 10.3× faster.
  • Select all 10,000 rows: 7.012 ms → 7.063 ms, essentially unchanged.

Outputs matched in each benchmark case. These measurements exclude file reads and decompression, so they describe assembly savings rather than end-to-end speedup. No actionable regressions found.

@platypii
platypii merged commit a51a523 into master Sep 17, 2026
6 checks passed
@platypii
platypii deleted the codex/bounded-variant-decoding branch September 17, 2026 22:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants