Skip to content

[BUG] Custom urltest latency checks depend on "Automatically remove servers that fail latency tests" checkbox #1967

Description

@JohnDoeWM

操作系统/Operating System

Windows

系统版本/Operating System Version

Windows 10

App版本/App Version

v1.2.26.2900

描述/Description

Related: #1919, #1909

The checkbox "Automatically remove servers that fail latency tests" controls more than just server removal - it gates the entire latency pipeline. Without it, latency tests may run (values appear in Zashboard), but the data is never applied, never persists, and the core never reloads.

Test results

Checkbox ON on stuck subscription:

  1. Enabled checkbox on the stuck subscription (the one with the dead node)
  2. Manually triggered latency test
  3. Core restarted, removed dead servers, connected to fastest node
  4. Other subscriptions (without checkbox) also started running latency tests over time

Checkbox ON on different subscription:

  1. Enabled checkbox on a different subscription (not the stuck one)
  2. Manually triggered latency test
  3. That subscription's dead servers were removed, core reloaded (same "reload after deleted failed" notification)
  4. Stuck subscription remained stuck, no other subscriptions ran latency tests

Checkbox OFF everywhere:

  1. Manually triggered latency test on stuck subscription
  2. Latency values appeared in Zashboard, but group remained stuck on dead node
  3. Restarted Karing - all latency data gone, still stuck
  4. Waited 10+ minutes - no subscription ran latency tests

Comparison

Checkbox ON (stuck) Checkbox ON (different) Checkbox OFF
Manual test runs Yes Yes Yes
Latency data appears Yes Yes (temporarily) Yes (temporarily)
Flagged subscription's dead servers removed Yes Yes No
Stuck subscription switches to fastest Yes No No
Data persists after restart Yes No No
Other subscriptions run latency tests Yes No No

Same groups, same rules, same subscriptions, same backup state, same health check interval (2 min). Only difference was which subscription had the checkbox.

Additional observations

  • Cascading failure: Disabling the stuck server doesn't help - the next one in line (same subscription) is also dead. Same 0/0 stall.
  • Manual removal doesn't help: Deleting dead servers by hand doesn't restore traffic. Only the checkbox triggers the core reload (_onEventLatencyUpdate → setServerAndReload()) that actually unblocks connection.
  • Ghost ping: A few nodes from other subscriptions show ping in Zashboard and change on restart, but carry 0/0 traffic. When the checkbox works correctly, ~10x more servers are pinged. Latency data is collected but never applied to selection.

Possible root cause in source code

In server_manager.dart schedulerTestLatency() (L1375), after all tests complete, removeLatencyError() is only called for subscriptions where testLatencyAutoRemove == true (L1434-1444). Only if change == true does the subscription get added to _latencyUpdatedConfigs, and only if this set is non-empty does the _onEventLatencyUpdate callback fire.

In home_screen.dart _onEventLatencyUpdate() (L1246), there is a second check: testLatencyAutoRemove must be true for at least one group. Only then does setServerAndReload() trigger a full core restart.

The retry count also depends on the flag (L1537):

int tryTimes = item.testLatencyAutoRemove ? 3 : 1;

testLatencyAutoRemove must be set on the subscription that owns the failing servers. It cannot be satisfied by any other subscription's flag.

This also explains ghost ping from other subscriptions - the core runs tests, but _latencyUpdatedConfigs stays empty because removeLatencyError() is never called with change == true for those subscriptions.

复现步骤/Reproduction steps

  1. Create Custom Auto Select groups across multiple subscriptions
  2. Reference them directly in rules (Rule mode)
  3. Set Health check interval to 2 minutes
  4. Connect - let a node go down
  5. Without checkbox, manually trigger latency test on stuck subscription
  6. Latency values appear in Zashboard, but no node switch - still stuck
  7. Restart Karing - latency data gone, still stuck
  8. Wait 10+ minutes - no subscription runs latency tests
  9. Enable checkbox on stuck subscription
  10. Manually trigger latency test
  11. Core restarts, dead servers removed, connected to fastest node
  12. Other subscriptions (without checkbox) also run latency tests over time

日志/Log

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions