Repository navigation
fix(serve): cap on-demand renders per /search request - #164
Merged
ASuresh0524 merged 1 commit intoOct 1, 2026
Merged
Conversation
With render-on-demand enabled (--kiwix-url), every hit whose tile is not on disk triggers a Chrome page render inside OnDemandTiles.chunk_path — serialized on a process-global lock and bounded only by PIXELRAG_RENDER_TIMEOUT (120s by default). Nothing capped how many one request could ask for. At this endpoint's own limits (32 queries from StarTrail-org#154, n_docs <= 1000 from StarTrail-org#122) a single include_images request could queue tens of thousands of renders, and because the lock is process-wide it stalls every other caller's renders too. No auth stands in front of it. Budget the renders per request (16). Past it a hit is returned without its image, which is already what a failed render or a missing tile yields, and the client can still fetch it from /tile. Cache hits must not spend the budget — in a warm deployment most calls are hits and cost only a stat. OnDemandTiles grows `cached_chunk_path`, which reports an already-rendered chunk without rendering, so the handler can tell the two apart before committing; `chunk_path` now routes through it, with the filename built in one place (`_chunk_file`). Not addressed: the global lock still serializes renders across requests, so a queue of modest requests can still pile up. Bounding queue depth or making renders concurrent is a larger design change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
@dex0shubham is attempting to deploy a commit to the andylizf's projects Team on Vercel. A member of the Team first needs to authorize it. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With render-on-demand enabled (
--kiwix-url),/searchhas no bound on howmuch work one request can queue. In the hit loop:
For every hit whose tile isn't on disk that reaches
OnDemandTiles.chunk_path→_render_and_chunk, which launches a Chrome pagerender, serialized on a process-global
_render_lockand bounded only byPIXELRAG_RENDER_TIMEOUT(120s by default).Nothing capped how many of those a single request may ask for. At this
endpoint's own limits — 32 queries (#154) and
n_docsup to 1000 (#122) — one{"include_images": true}request can queue tens of thousands of serializedrenders. The handler
awaits them one at a time, so the request occupies a taskfor as long as that takes, and because the lock is process-wide, every other
caller's renders stall behind it. The endpoint is public and unauthenticated.
This is the same family as the
n_docsand query-image bounds, but a bound ontime rather than memory — the amplification those two didn't close.
Fix
A per-request render budget (16). Past it the hit is still returned, just
without
image_base64— already what a failed render or a missing tile yields,so no new client contract — and the client can fetch it from
/tile/{article_id}/{tile_index}/{chunk_index}. A real caller asks forn_docs=10, so a legitimate request never reaches the cap.Cache hits must not spend the budget. In a warm deployment most calls are
hits that cost a single
stat, and charging them would withhold images fromalready-rendered articles for no reason.
OnDemandTilesgrowscached_chunk_path, which reports an already-rendered chunk withoutrendering, so the handler can tell a hit from a render before committing to one.
chunk_pathnow routes through it, with the chunk filename built in one place(
_chunk_file) instead of two.Tests
tests/test_serve_ondemand_budget.pydrives the real handler throughTestClientwith pre-computed query embeddings — no model, no index, no Chrome —against a counting stand-in for
OnDemandTiles:Removing the budget makes them fail with the amplification itself:
40 != 16for one query,
160 != 16for four — and that is atn_docs=40, far below whatthe endpoint permits.
Full suite: 165 passed, 1 skipped (
--extra serve --extra qdrant).Not addressed
The global lock still serializes renders across requests, so a queue of
modest requests can pile up even with each one budgeted. Bounding queue depth,
or making renders concurrent, is a larger design change and I'd rather not
half-build it here — happy to follow up if you'd like a particular shape.
Also left alone:
_MAX_ONDEMAND_RENDERSis a module constant, not an env knoblike
PIXELRAG_RENDER_TIMEOUT. Easy to change if you'd prefer operators tune it.