Skip to content

[cleanup_openstack] Force-drain stuck OpenStack namespace - #4241

Open
mnietoji wants to merge 1 commit into
openstack-k8s-operators:mainfrom
mnietoji:fix-cleanup-stuck-openstack-namespace
Open

mnietoji wants to merge 1 commit into
openstack-k8s-operators:mainfrom
mnietoji:fix-cleanup-stuck-openstack-namespace

Conversation

@mnietoji

@mnietoji mnietoji commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Teardown could fail with the openstack namespace stuck in Terminating forever. wait_and_cleanup_resource.yaml only force-removes finalizers from the top-level CRs (OpenStackControlPlane, dataplane, RabbitmqCluster), but their children (keystone*, nova*, galera, mariadbaccount, keystoneendpoint, ...) and core resources that carry OpenStack finalizers (secrets, configmaps, services) can be orphaned with dangling finalizers once the operators are removed. The namespace deletion in Phase 2 then times out with "Namespace" "openstack": Timed out waiting on resource.

Wrap the Phase 2 operator/namespace removal in a block/rescue: on failure, force_drain_namespace.yaml discovers every namespaced *.openstack.org (and rabbitmq.com) CRD plus the relevant core kinds and strips the finalizers from whatever is left, then the namespace deletion is retried. The operator-removal wait is shortened (cifmw_cleanup_openstack_operators_wait_timeout, default 300s) so a stuck namespace is recovered quickly instead of burning the full default timeout.

@qodo-code-review

Copy link
Copy Markdown

Qodo reviews are paused for this user.

Troubleshooting steps vary by plan Learn more →

On a Teams plan?
Reviews resume once this user has a paid seat and their Git account is linked in Qodo.
Link Git account →

Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center?
These require an Enterprise plan - Contact us
Contact us →

@openshift-ci

openshift-ci Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign bshewale for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@mnietoji
mnietoji force-pushed the fix-cleanup-stuck-openstack-namespace branch 3 times, most recently from 11c2e12 to 8cfac6b Compare October 8, 2026 09:31
@mnietoji

mnietoji commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

/retest

@mnietoji
mnietoji force-pushed the fix-cleanup-stuck-openstack-namespace branch from 8cfac6b to e5f8b56 Compare October 8, 2026 09:43
Teardown could fail with the `openstack` namespace stuck in `Terminating`
forever. `wait_and_cleanup_resource.yaml` only force-removes finalizers from
the top-level CRs (`OpenStackControlPlane`, dataplane, `RabbitmqCluster`),
but their children (`keystone*`, `nova*`, `galera`, `mariadbaccount`,
`keystoneendpoint`, ...) and core resources that carry OpenStack finalizers
(secrets, configmaps, services) can be orphaned with dangling finalizers once
the operators are removed. The namespace deletion then times out with
`"Namespace" "openstack": Timed out waiting on resource`.

Make the force-drain a true safety net. Phase 2 (remove operators + delete
the `openstack` namespace) runs in a `block`; issuing the namespace deletion
already stamps a `deletionTimestamp` on everything inside it. If that times
out because leftover CRs are wedged on finalizers whose operators are gone,
the `rescue` runs `force_drain_namespace.yaml` to discover every namespaced
`*.openstack.org` (and `rabbitmq.com`) CRD plus the relevant core kinds, strip
their finalizers, and retry the removal. The wait is kept short via
`cifmw_cleanup_openstack_operators_wait_timeout` so recovery kicks in quickly
instead of burning the full default timeout. On a clean teardown the rescue is
never reached, so the happy path pays zero extra cost.

The duplicated finalizer-stripping logic is unified into
`strip_finalizers.yaml`, reused both by the per-kind `rescue` in
`wait_and_cleanup_resource.yaml` and by `force_drain_namespace.yaml`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Miguel Angel Nieto Jimenez <mnietoji@redhat.com>
@mnietoji
mnietoji force-pushed the fix-cleanup-stuck-openstack-namespace branch from e5f8b56 to 1bbf9c4 Compare October 8, 2026 09:49
@centosinfra-prod-github-app

Copy link
Copy Markdown

Build failed (check pipeline). Post recheck (without leading slash)
to rerun all jobs. Make sure the failure cause has been resolved before
you rerun jobs.

https://gateway-cloud-softwarefactory.apps.ocp.cloud.ci.centos.org/zuul/t/rdoproject.org/buildset/3cedaf4dff0f42468e54c730559d7cbb

✔️ openstack-k8s-operators-content-provider SUCCESS in 2h 36m 16s
✔️ podified-multinode-edpm-deployment-crc SUCCESS in 1h 19m 02s
✔️ podified-multinode-edpm-deployment-crc-centos-10 SUCCESS in 1h 17m 18s
✔️ cifmw-crc-podified-edpm-baremetal SUCCESS in 1h 30m 10s
✔️ cifmw-crc-podified-edpm-baremetal-minor-update SUCCESS in 1h 59m 33s
✔️ cifmw-pod-zuul-files SUCCESS in 4m 10s
✔️ openstack-k8s-operators-content-provider-bootc SUCCESS in 1h 10m 03s
❌ cifmw-crc-podified-edpm-baremetal-bootc NODE_FAILURE Node(set) request 099-0000228627 failed in 0s
✔️ noop SUCCESS in 0s
✔️ cifmw-pod-ansible-test SUCCESS in 8m 47s
✔️ cifmw-pod-pre-commit SUCCESS in 8m 20s
✔️ cifmw-molecule-cleanup_openstack SUCCESS in 2m 18s

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant