Skip to content

Reuse the element scratch buffer across rows in json_get_array - #142

Merged
adriangb merged 2 commits into
mainfrom
reuse-json-get-array-scratch
Sep 23, 2026
Merged

adriangb merged 2 commits into
mainfrom
reuse-json-get-array-scratch

Conversation

@samuelcolvin

@samuelcolvin samuelcolvin commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-up to #139.

The json_get_array fast path buffered each row's element slices in a SmallVec<[&str; 8]>, which allocates once per row for arrays longer than eight elements.
This replaces it with a Vec<Range<usize>> created once per batch and cleared per row, so its capacity is reused across rows. The buffer grows only when a row has more elements than any earlier row in the batch, so the allocation count per batch depends on the widest row, not on the number of rows.
Storing byte ranges rather than borrowed &str is what lets one buffer outlive each per-row closure call.

Elements are now sliced from the input &str by range, an O(1) char-boundary check, instead of a std::str::from_utf8 pass over each element.
jiter stops on ASCII structural bytes, so every range boundary is a char boundary.

invoke_array_scalars_direct takes FnMut so the closure can capture the buffer.
The smallvec dependency is removed.

Null-on-error behaviour is unchanged: ranges are only sliced after the whole array parses, and the existing malformed_late_element_does_not_append_partial_values test still covers that.
The equivalence test runs rows of 5, 9, 2, 0 and 1 elements in sequence through the shared buffer.

Benchmarks

The existing json_get_array_array bench uses eight elements, which fit the old inline buffer, so this adds json_get_array_array_wide with 32 elements.
Both are 1,024-row Utf8View batches, cargo bench --bench main release builds on an Apple M3 Max, criterion --save-baseline on main's source and --baseline on this branch, same bench file for both.
Two before/after pairs, reported as criterion's mean estimate:

bench elements before after change
json_get_array_array 8 382 µs 331 µs -12.8%
json_get_array_array 8 373 µs 322 µs -13.0%
json_get_array_array_wide 32 1.490 ms 1.200 ms -18.0%
json_get_array_array_wide 32 1.485 ms 1.179 ms -20.0%

The eight-element gain comes from dropping the per-element from_utf8 pass; the extra gain at 32 elements is the per-row heap allocation the old buffer spilled into.
A first baseline run measured 487 µs and 1.876 ms, well outside the other two, and was discarded as noise.

cargo test and cargo clippy --all-targets -- -D warnings pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_01L2MYrU7T15KLX6ASKAMu7Y

The fast path added in #139 collected each row's element slices in a
SmallVec<[&str; 8]>, which allocates once per row for arrays longer than
eight elements. The buffer is now a Vec<Range<usize>> created once per
batch and cleared per row, so it allocates at most once per batch
regardless of array length. Storing byte ranges instead of borrowed
slices is what lets one buffer outlive the per-row closure call.

The row's elements are sliced from the input &str by range, which is an
O(1) char-boundary check, in place of the previous std::str::from_utf8
pass over each element. jiter stops on ASCII structural bytes, so the
boundaries are always valid.

invoke_array_scalars_direct takes FnMut so the closure can capture the
buffer. The smallvec dependency is removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L2MYrU7T15KLX6ASKAMu7Y
@codecov-commenter

codecov-commenter commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.83333% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 93.09%. Comparing base (2538b69) to head (c607e24).

Files with missing lines Patch % Lines
src/json_get_array.rs 95.45% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #142      +/-   ##
==========================================
+ Coverage   93.05%   93.09%   +0.03%     
==========================================
  Files          18       18              
  Lines        1815     1824       +9     
  Branches     1815     1824       +9     
==========================================
+ Hits         1689     1698       +9     
  Misses         70       70              
  Partials       56       56              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The existing json_get_array_array bench uses eight elements, which never
spilled the SmallVec inline buffer, so it does not exercise the case the
scratch-buffer reuse targets.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L2MYrU7T15KLX6ASKAMu7Y

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation preserves behavior and is adequately tested, with only a minor allocation-claim discrepancy.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
What changed in this PR

Reuses a batch-scoped scratch buffer in json_get_array to reduce parsing allocations.

Changes:

  • Stores element byte ranges in a reusable Vec.
  • Allows mutable direct-path callbacks.
  • Removes smallvec and adds a wide-array benchmark.
File Description
src/​json_get_array.rs Implements reusable range buffering.
src/​common.rs Accepts FnMut callbacks.
Cargo.toml Removes smallvec.
benches/​main.rs Adds a 32-element benchmark.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/json_get_array.rs
@adriangb
adriangb merged commit 5cfae72 into main Sep 23, 2026
8 checks passed
@adriangb
adriangb deleted the reuse-json-get-array-scratch branch September 23, 2026 14:03
@adriangb adriangb mentioned this pull request Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants