Skip to content

SFDP: Possible worker deadlock when killing sessions #3731

Description

@vrpolakatcisco

I am not yet sure how frequent this is in rls2606. CSIT default configuration was not capturing the core, but after some tweaks I got a verify run. In NDRPDR CPS test there are multiple trials [0] with sfdp_kill_session(is_all=True) PAPI call in between, plus some debug CLI outputs I am not sure are important. The core is here [1] and I think it shows some CLI command (subsequent debug output) timing out when worker is clearing the session table. There is no good reason why table cleanup should take ~1s to trigger the timeout, but I am not sure what exactly is wrong.

[0] https://logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-verify-2n-grc/30269880859/log.html.gz#s1-s1-s1-s1-s2-t1-k2-k14-k16
[1] https://logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-verify-2n-grc/30269880859/log.html.gz#s1-s1-s1-s1-s2-t1-k3-k4-k1

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions