Repository navigation
fix: complete recoverable file cleanup (2B) - #2846
Conversation
rogercloud
left a comment
There was a problem hiding this comment.
Approach: acceptable-with-reservations — the per-file execution lock sits on hot read paths as a blocking exclusive lock (first two Major items); the claim/phase-receipt/settlement model itself fits the existing compensation lifecycle.
Major
src/xagent/web/services/uploaded_file_cleanup_publication.py:105— blocking 15sfile_cleanup_lockis taken on everyensure_local/materialize(even with a local copy) and called directly on the event loop from async handlers insrc/xagent/web/api/files.py(download/preview/public_*);_convert_pptx_to_pdfholds the same lock across soffice. Fix: return an existing local copy lock-free, lock only to publish, useasyncio.to_thread, and narrow the PPTX lock to validate+publish.
Trigger: PPTX preview-pdf in progress + any download/preview of the same file on one worker. Impact: loop blocks up to 15s, converter stalls, second request gets 503, all requests on that worker stall.src/xagent/web/services/uploaded_file_cleanup_resources.py:164—uploadsis resolved butpath(from the unresolved upload path) is not, so with a symlinked uploads diris_relative_to(owner_root)is False and the source is only "preserved". Fix: canonicalize realpath(parent)/name before containment and allowlist checks; add a symlinked-root test.
Trigger: uploads dir behind a symlink. Impact: row settled "deleted" and removed while the local upload stays on disk with no handle.src/xagent/web/services/uploaded_file_cleanup_resources.py:102— nested temp/quarantine names (.{name}.cleanup-{32hex}= N+42; restore temp inmanaged_file_ref.py+storage.py= N+53 vs N+28 on base) exceed 255 bytes for long names. Fix: bounded names from a hash of file_id/name plus a random suffix; add a 250-byte filename test.
Trigger: uploads with ~203+ byte names (about 68+ CJK chars). Impact: restore returns 503; names >=214 bytes can never be quarantined and stay "pending" after the durable object is deleted.
Minor
src/xagent/web/services/uploaded_file_store.py:1921— delete CAS uses the full version snapshot (broader than a narrow guard) and KB delete callers (kb_collection_service.py,kb_file_service.py) neither skip compensating rows nor catchUploadedFileVersionConflict(500 on collection delete). Thewith_for_update()row lock is also held acrossdelete_durable()/unlink whenafter_commitis None, blocking a concurrent claim on PostgreSQL. Fix: CAS on id/file_id/user_id + status != compensating, callers skip compensating rows, lock only the final conditional delete.uploaded_file_cleanup.py:111-121— durable phase runs before already-recorded uncertainty (missing SHA-256, symlink source) is evaluated, leaving a compensating row that recovery re-warns on forever. Fix: fail or skip the claim before the durable phase, or persist a needs_reconciliation marker.uploaded_file_cleanup_publication.py:48-75— claimed/fenced/version-mismatch raiseDurableObjectMissingError, whichfiles.pyhandlers treat as "servefile_ref.local_path". Fix: distinct exception mapped to 404.uploaded_file_cleanup_publication.py:94-101— guard callsrelease_db_connection_if_cleanon the caller's session and raises bareRuntimeErrorfor a dirty one. Fix: don't release a session the adapter doesn't own; raiseDurableStorageOperationError.uploaded_file_cleanup_resources.py:196-225—os.listdirof the source parent per manifest is O(N*M) for flatuser_{id}/dirs at 500 per GC page. Fix: scandir with prefix match or per-directory cache.uploaded_file_cleanup_resources.py/storage.py— layout (materialize dirs, preview ownership, privatestorage._scoped) is re-derived in cleanup; a layout change silently leaks (S3 objects withoutxagent-sha256metadata already leak cache copies). Fix: export path/matcher helpers from storage/preview layers.file_reference.py:69-78—file_cleanup_lockon every cache hit creates a never-removed<sha>.cleanup.lockper file read and serializes readers; docs frame growth as cleanup-only. Fix: lock-free fast path, recheck under lock when publishing, document accurately.- Tests: private filelock internals patched (
FileLock._acquire,_context.lock_file_fd) whilefilelock>=3.0.0is allowed; in-flight recovery test dropped the "exists keeps compensating, later absent settles" sequence; "external" case never exercisesXAGENT_EXTERNAL_UPLOAD_DIRS; no symlinked-root test; pptx contention case stubs_resolve_file_pathwithfile_record=None, so the asserted 503 never comes from the materialize guard.
Simplification
src/xagent/web/api/files.pyL497: drop thewritten_identitiesNone default (single caller passes it; guard at 517-518 is dead). Make it required.src/xagent/web/services/uploaded_file_cleanup.pyL86/127/147: three_save_manifestcalls repeat 7 identity args. Bind once withfunctools.partial.src/xagent/web/services/uploaded_file_cleanup_resources.pyL198-217: root-equality is evaluated per entry. Hoistprefixesonce.- net: -19 lines possible
Blocking: yes — recommended event: REQUEST_CHANGES
src/xagent/web/services/uploaded_file_cleanup_publication.py:105, Major, blocking lock on event loop stalls the worker and 503s concurrent preview/download, [new]src/xagent/web/services/uploaded_file_cleanup_resources.py:164, Major, symlinked uploads dir leaves local file orphaned while row is deleted, [new]src/xagent/web/services/uploaded_file_cleanup_resources.py:102, Major, long filenames hit ENAMETOOLONG on restore and can never be quarantined, [new]
Keep managed cache reads responsive, retain exact publication generations, and bound recoverable resource names. Defer known uncertainty before destructive storage work and preserve KB deletion races. Refs xorbitsai#1086
|
Follow-up to review 5426049616. Addressed this round in
The PPTX cache-miss producer retains its guard while its subprocess can write owned temporary output. Locking only the final rename would let cleanup settle and remove its recovery handle before the converter recreates output. Cached readers no longer wait and the event loop stays responsive; narrowing the remaining producer interval would require a separate producer/recovery contract. Immediate direct-deletion SQL-lock scope is tracked in #2848, and existing preview locator/matcher consolidation in #2849. Per-manifest discovery still scales with directory size; batching/bounds remain #1086 2C work-budget scope. Retired fences and live lock inodes remain monotonic as required. The existing explicit Self-review found and fixed gaps in the initial revision: the intermediate 210-byte nested copy could lose its owner prefix; manifest capture briefly depended on provider hash reads; filename/backend/URI generation checks were incomplete; and previously NULL generation tokens were skipped. Production-call/process regressions cover all four. Verification: 169 focused SQLite/PostgreSQL cases passed; 473 affected regressions passed with two KB-ingest-unavailable skips; Ruff, pinned isort/codespell, whitespace, and focused nine-file mypy passed. The full suite stopped at three pre-existing Alembic log assertions, reproduced on the exact reviewed head with the same environment (three failures, 298 passes). Parent-commit regressions fail on the actual API, collector, restoration, KB deletion and after-commit paths. This completes the fixes for this review round, not issue #1086 or its remaining 2C/legacy delivery. The approved Standards/Spec preflight is open: no blocker; two nonblocking readability/field-list observations. Final configured pre-commit also passed, including its isolated mypy hook. Please re-review the current head. |
rogercloud
left a comment
There was a problem hiding this comment.
Minor
src/xagent/web/services/uploaded_file_cleanup_publication.py:149: the lock-free fast path skips the existing materialize cache, so concurrent reads return 503 during PPTX conversion (see inline comment).src/xagent/web/api/chat.py:349,src/xagent/core/tools/adapters/vibe/file_ingestion_tool.py:189,src/xagent/web/websocket.py:1362(viaresolve_turn_file_infos): async callers still run the guarded syncensure_localon the event loop. Each call now runs a SQL validation, and on a cache miss with same-file contention it can wait up to 15s onFileLock, which blocks the loop. Offload these calls withasyncio.to_threadorasync_managed_copy.src/xagent/web/services/uploaded_file_cleanup_publication.py:44: callers don't mapFilePublicationUnavailableto missing-file handling, so a claim race returns 500 or fails the run (see inline comment).src/xagent/web/services/uploaded_file_cleanup_publication.py:59: snapshot-backed refs lose the row id, so the row-id check is skipped (see inline comment).src/xagent/web/services/uploaded_file_cleanup.py:108: an uncertain manifest leaves the claim committed, so every recovery sweep retries it and counts it as failed (see inline comment).src/xagent/web/api/files.py:694: one DB error in the finalizer aborts cleanup of the remaining staged uploads (see inline comment).tests/web/services/test_uploaded_file_recovery.py:294: the test no longer covers recovery deferring onexistsand then settling onabsent(see inline comment).tests/web/services/test_complete_upload_cleanup.py:1264: the PPTX contention test never reaches the lock guard (see inline comment).
Simplification
src/xagent/web/services/uploaded_file_cleanup_resources.pyL55: delete:_quarantine_namestill has a readable.{name}.cleanup-…branch next to the hashed one. No released format wrote that form, and the manifest stores whichever name was chosen. Always return the hashed.cleanup-{sha24}-{gen16}form.src/xagent/web/services/uploaded_file_cleanup_publication.pyL84: shrink:_prepare_copyreturns(clone, sessions, read_only), but both extras are already stored on the clone andasync_managed_copyignores them. Return only the clone.src/xagent/web/services/uploaded_file_cleanup_publication.pyL37: delete:"task_id"in_COPY_FIELDSis never read from the snapshot. Drop it.
net: -7 lines possible
Blocking: no — recommended event: APPROVE
Map claimed or retired uploads to existing missing-file behavior, offload async copy waits, and retain complete snapshot generation identity. Read validated materialized caches before the execution lock, isolate staged-path cleanup failures, and strengthen production contention regressions. Track retained cleanup uncertainty in xorbitsai#2851. Refs xorbitsai#1086.
|
Addressed this review round in one commit: 2b2018be.
Retained uncertainty classification/accounting is explicitly deferred to P2 #2851; conservative retention stays intact. Validation: 288 focused SQLite/PostgreSQL and caller tests passed; 316 additional affected regressions passed; all configured pre-commit hooks passed, including mypy across 880 source files. Full-suite execution encounters three Alembic INFO-log assertions, reproduced on the exact parent Deployment requirements and the existing 2C/legacy/fence-compaction dependencies remain documented. This does not complete #1086. Refs #1086. |
Problem and behavior
Refs #1086.
The task-less collector could unlink a local upload before winning its SQL claim. Registered compensation and recovery could delete the upload row after durable-object absence while preview or local cleanup was still unfinished, losing their recovery handle.
This implements 2B complete, recoverable cleanup on the existing compensation lifecycle. An exact claim commits before destructive I/O. A focused JSON manifest retains resource locators, configuration, ownership/replacement evidence, quarantine names, and durable/local/preview phase receipts. The upload row remains unavailable and recoverable until every required phase succeeds and exact-generation settlement commits. Already absent resources succeed; uncertain outcomes remain pending.
A separate persistent execution lock serializes takeover, storage work, settlement, managed-copy restoration, and owned preview publication. Reference locks and SQL connections are released before slow storage work. Sources require managed owner-root containment and regular-file/inode/checksum evidence; external/shared sources are preserved. Owned partial copies and converter temporaries are included. Replacements and escaping symlinks remain identifiable for reconciliation. Retired SQL identities and standalone fences survive upload-row removal.
Request cancellation and existing compensation errors remain primary. Request finalization preserves protected/rebound uploads, unfinished claims, and replacement inodes. Direct deletion receives only a narrow guard against erasing a compensating row.
Verification
/previewresponses, claimed/retired file handling in Chat/WebSocket/agent registration, KB event-loop responsiveness, snapshot row removal and NULL-generation changes, and staged-path cleanup after a transient ownership-query failure. Fourteen new regression cases fail on the exact reviewed parent951aa7b4; the corrected head passes them.DurableStorageOperationError->filelock.Timeoutcause chain, and proof that no converter subprocess starts.951aa7b4reproduces 3 failed, 298 passed in an independent worktree with the same environment. The two affected Alembic files pass in isolation (40 tests), establishing collection/logging-order sensitivity rather than a cleanup regression.Standards/Spec review found no blocker; the only judgement call was non-blocking duplication between the explicit sync/async turn-file result loops. The user explicitly chose to skip this round's pr-preflight before pushing
951aa7b4..2b2018be.Caller and publication contracts
Available local copies and fresh preview caches use existing read paths. Materialized cache reads validate scoped storage routing/checksum and current persistent row generation before and after probing, without taking the execution guard. Corrupt cache repair still waits for the guard.
ensure_localcontinues restoring its original target; dirty transactions remain intact and cannot probe storage or publish bytes.Chat, WebSocket attachment resolution and KB ingestion snapshot metadata on the owning thread, return clean request connections and await detached storage work off-loop. Claimed/retired uploads map to each caller's existing missing-file behavior without exposing a claimed local path. Turn, KB and version snapshots retain persistent row identity and complete generation tokens, including nullable backend/URI/ETag changes. Publication lock contention remains a typed storage/503 boundary rather than an unhandled request error.
Uncached preview conversion holds its execution guard while owned temporary output can be created, drains work on cancellation and releases the guard before propagating it. Request finalization isolates ownership-query failures per staged path, preserves uncertain ownership/replacement evidence and continues cleaning the rest while keeping the primary error/cancellation contract. Configured-root aliases, external roots, bounded UTF-8 names and provider-specific cache namespaces retain their ownership gates. Deferred KB deletion publishes storage cleanup only after its exact generation CAS succeeds and the caller commits.
Deployment
Stop upload/local/preview producers, KB reference writers, collectors, and compensation workers; apply online migration
20261005_uploaded_file_cleanup_manifest; restart together on the same version. Preserve the database, resource roots, and shared LanceDB coordination directory. Mixed-version workers cannot honor publication/execution fencing. Downgrade refuses pending manifests; complete or reconcile those claims first.See complete cleanup design and interruption protocol for phases, lock order, matching axes, and deployment constraints. This single combined PR exceeds the opening size/footprint budget with explicit user authorization.
Remaining delivery and known out-of-scope defects
This does not finish #1086. No detached collector, seven-day TTL, scan indexes, scheduling, or work budget is enabled; those remain 2C. Historical inventory and legacy/local-only backfill remain later work. Unknown historical resources are preserved for reconciliation.
SQL cleanup fences,
.claimedmarkers, reference.lockfiles, and execution.cleanup.lockfiles remain monotonic. Safe compaction needs separate work; retired identities are not erased and live lock inodes are never replaced to reduce growth.