You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Tracking issue for the watchdog stall first observed during benchmarking of PR #5799 (now superseded by PR #5819) and not yet diagnosed. The same payload is carried forward in #5819, so until this is closed out it remains an open risk for that PR.
Observed symptom
During a sustained 500 client / 50 pool benchmark (4 KB rows, SSL on, 120 s), a watchdog abort fired post-benchmark with the message:
Per-group naming in src/main.cpp:2674,2714 — the %s is literally \"MySQL\" or \"PostgreSQL\". Worth confirming on re-run whether it was actually "MySQL" or paraphrased — the benchmark itself was PgSQL.
Watchdog mechanics (for reference)
Heartbeat source
Each worker writes atomic_curtime = curtime after every poll() return — MySQL_Thread.cpp:3723 / PgSQL_Thread.cpp:3269
Watchdog loop interval
200 ms — src/main.cpp:2586
"Missed" threshold
curtime > atomic_curtime + (poll_timeout_ms + 1000) × 1000 → 3000 ms with default 2000 ms poll_timeout — src/main.cpp:2550
Self-starvation auto-reset
If the watchdog itself was scheduled too slowly (>300 ms × inner_loops), it resets all groups — src/main.cpp:2608
Abort threshold
10 consecutive missed checks → assert(0)
For an abort, a worker must fail to refresh atomic_curtime for >3 s, ten checks in a row. That's a real stall — the watchdog auto-immunizes against its own scheduling delays, so this isn't a measurement glitch.
Before: 1 batched wrlock() per outer iteration (via push_MyConn_to_pool_array).
After: 1 individual wrlock() per release for (N−1)/N of releases.
End-of-benchmark client-disconnect flood now produces hundreds of individual writer-locks per thread in rapid succession. If a Monitor or Admin wrlock holder overlaps, all workers serialize on the writer-lock and stop updating atomic_curtime. Three seconds of contention is plausible at thousands of releases/sec colliding with a long Monitor critical section.
2. Stale pause_until bug from #5784 is NOT in this stack (still present at lib/Base_Thread.cpp:403). On its own it only causes a single 2 s poll() block — under the 3 s threshold — but stacked behind other contention it widens the window.
3. Not plausible as direct causes: partition gate state machine, partition Pass 1, ERR_clear_error placement fix, -fno-omit-frame-pointer flag. None of these can produce multi-second blocks.
Capture perf record -F 99 -g --call-graph fp -p <pid> for ~30 s around the post-benchmark moment, plus a pstack of all threads at the abort. With this stack's -fno-omit-frame-pointer build flag, call graphs are clean. pstack will show immediately whether workers are blocked in pthread_rwlock_wrlock (HGM), pthread_mutex_lock (thread_mutex), poll, or elsewhere.
If lock contention on MyHGM->wrlock() / PgHGM->wrlock() is confirmed, the fix is bounded batching: cache locally and flush via push_MyConn_to_pool_array whenever the local cache exceeds a threshold (e.g., max(1, conns_per_thread / N)). Same fairness, batched lock cost.
Tracking issue for the watchdog stall first observed during benchmarking of PR #5799 (now superseded by PR #5819) and not yet diagnosed. The same payload is carried forward in #5819, so until this is closed out it remains an open risk for that PR.
Observed symptom
During a sustained 500 client / 50 pool benchmark (4 KB rows, SSL on, 120 s), a watchdog abort fired post-benchmark with the message:
(Source: PR #5799 description.)
Per-group naming in
src/main.cpp:2674,2714— the%sis literally\"MySQL\"or\"PostgreSQL\". Worth confirming on re-run whether it was actually "MySQL" or paraphrased — the benchmark itself was PgSQL.Watchdog mechanics (for reference)
atomic_curtime = curtimeafter everypoll()return —MySQL_Thread.cpp:3723/PgSQL_Thread.cpp:3269src/main.cpp:2586curtime > atomic_curtime + (poll_timeout_ms + 1000) × 1000→ 3000 ms with default 2000 mspoll_timeout—src/main.cpp:2550src/main.cpp:2608assert(0)For an abort, a worker must fail to refresh
atomic_curtimefor >3 s, ten checks in a row. That's a real stall — the watchdog auto-immunizes against its own scheduling delays, so this isn't a measurement glitch.Candidate causes ranked by likelihood
1. 1-in-N cache change (most likely) —
lib/MySQL_Thread.cpp:6478,lib/PgSQL_Thread.cpp:5833.Inverts the lock pattern:
wrlock()per outer iteration (viapush_MyConn_to_pool_array).wrlock()per release for (N−1)/N of releases.End-of-benchmark client-disconnect flood now produces hundreds of individual writer-locks per thread in rapid succession. If a
MonitororAdminwrlock holder overlaps, all workers serialize on the writer-lock and stop updatingatomic_curtime. Three seconds of contention is plausible at thousands of releases/sec colliding with a long Monitor critical section.2. Stale
pause_untilbug from #5784 is NOT in this stack (still present atlib/Base_Thread.cpp:403). On its own it only causes a single 2 spoll()block — under the 3 s threshold — but stacked behind other contention it widens the window.3. Not plausible as direct causes: partition gate state machine, partition Pass 1,
ERR_clear_errorplacement fix,-fno-omit-frame-pointerflag. None of these can produce multi-second blocks.Suggested investigation
Minimum reproduction + diagnosis path:
perf record -F 99 -g --call-graph fp -p <pid>for ~30 s around the post-benchmark moment, plus apstackof all threads at the abort. With this stack's-fno-omit-frame-pointerbuild flag, call graphs are clean.pstackwill show immediately whether workers are blocked inpthread_rwlock_wrlock(HGM),pthread_mutex_lock(thread_mutex),poll, or elsewhere.MyHGM->wrlock()/PgHGM->wrlock()is confirmed, the fix is bounded batching: cache locally and flush viapush_MyConn_to_pool_arraywhenever the local cache exceeds a threshold (e.g.,max(1, conns_per_thread / N)). Same fairness, batched lock cost.Blocking relationship