From 60276039334fd9c9ffa52190aee08cc93e9fb587 Mon Sep 17 00:00:00 2001 From: Michael Agun Date: Tue, 11 Aug 2026 16:17:52 -0700 Subject: [PATCH 1/3] netebpfext: purge stale WFP state at init and fail loudly when it cannot PR #5472 guarantees that driver unload always completes, even when a WFP filter delete permanently fails. That exposes the next problem: the filter that could not be deleted is still installed in WFP. WFP management objects are not owned by the driver that created them, and this extension's engine sessions are non-dynamic, so closing the engine handle does not remove them. A surviving filter therefore pins the callout, sub-layer and provider objects it references, their deletion at unload fails too, and the next FwpmProviderAdd fails with FWP_E_ALREADY_EXISTS (0xC0220009). Observed on a test machine, with the service still reporting STATE: 4 RUNNING afterwards. A stale filter is also actively dangerous. Its rawContext still holds the address of a filter context that the previous driver instance allocated and freed. Today that is inert only because initialization fails at FwpmProviderAdd and so never reaches FwpsCalloutRegister. The moment a callout is registered under a GUID a stale filter references, WFP begins invoking classifyFn with a dangling pointer. Three changes, which have to travel together: * Add _net_ebpf_ext_purge_stale_wfp_objects(), called after FwpmEngineOpen and before FwpmTransactionBegin, in its own transaction so that aborting the initialization transaction cannot roll the purge back. It deletes objects tagged with EBPF_WFP_PROVIDER in reverse dependency order: filters, callout objects, sub-layers, then the provider. Running it before the first FwpsCalloutRegister is what makes it safe: with no callout function registered, WFP delivers no delete notification for the filters removed here, so the purge never touches a stale rawContext. Every *_NOT_FOUND result is success, so a clean boot is a handful of no-op calls. If any eBPF filter survives the purge, initialization fails. A partial purge followed by a successful start is strictly worse than not starting, because it is exactly the case that registers callouts alongside dangling contexts. * Check the result of net_ebpf_extension_initialize_wfp_components() in _net_ebpf_ext_driver_initialize_objects() instead of discarding it with (void). DriverEntry already calls _net_ebpf_ext_driver_uninitialize_objects() and returns the failing status, so the service start now fails with a real error code rather than reporting RUNNING with no WFP components registered. This is a deliberate behaviour change: netebpfext now refuses to load when WFP initialization fails, instead of loading in a silently inert state. Closes the TODO for issue #521. * Clear rundown_acquired once create_filter_context() succeeds, in _net_ebpf_extension_hook_provider_attach_client(). The provider rundown reference is acquired by the attach frame, but ownership transfers to the filter context the instant it exists; from then on the context releases it, either when its refcount reaches zero or via the unload sweep for a zombie. Without this, a failure between context creation and the end of the function would release the same reference twice. That path is currently unreachable, but it becomes a live rundown underflow as soon as one is added, and the purge work above adds failure paths to this file's neighbourhood. --- netebpfext/net_ebpf_ext.c | 315 ++++++++++++++++++++++++++++++ netebpfext/sys/net_ebpf_ext_drv.c | 33 +++- 2 files changed, 344 insertions(+), 4 deletions(-) diff --git a/netebpfext/net_ebpf_ext.c b/netebpfext/net_ebpf_ext.c index 6bbdc7acdb..7152235101 100644 --- a/netebpfext/net_ebpf_ext.c +++ b/netebpfext/net_ebpf_ext.c @@ -793,6 +793,308 @@ net_ebpf_ext_uninitialize_ndis_handles() } } +/** + * @brief Number of WFP objects requested per enumeration call during the stale-object purge. Purely a batching + * choice: the purge loops until the enumeration is drained, so this only trades call count against transient + * allocation size. + */ +#define NET_EBPF_EXT_PURGE_ENUM_BATCH_SIZE 64 + +/** + * @brief Upper bound on enumeration batches during the purge. A WFP enumerator is a snapshot and always advances, + * so this bound is never reached in practice; it exists so that a misbehaving engine cannot spin driver + * initialization forever. Reaching it is treated as a purge FAILURE, never as a completed purge: if the + * enumeration did not drain, we cannot prove that no stale filter remains. + */ +#define NET_EBPF_EXT_PURGE_MAX_ENUM_BATCHES 1024 + +/** + * @brief Deletes every WFP object tagged with the eBPF provider that survived a previous driver instance. + * + * A WFP filter whose delete permanently failed stays installed after the driver unloads: WFP management objects + * are not owned by the driver that created them, and this extension's engine sessions are non-dynamic, so closing + * the engine handle does not remove them. Such a filter pins the callout, sub-layer and provider objects it + * references, so their deletion at unload fails too, and the next FwpmProviderAdd fails with FWP_E_ALREADY_EXISTS. + * + * A stale filter is also actively dangerous. Its rawContext still holds the address of a filter context that the + * previous driver instance allocated and freed. Nothing revives that pointer, so the moment a callout is + * registered under a GUID the stale filter references, WFP starts invoking classifyFn with a dangling pointer. + * The caller must therefore treat failure here as fatal: starting with stale filters still installed is strictly + * worse than not starting at all. + * + * This is why the purge must run before any FwpsCalloutRegister call. With no callout function registered, WFP + * delivers no delete notification for the filters removed here, so the purge itself never touches the stale + * rawContext values. + * + * Objects are removed in reverse dependency order (filters, callouts, sub-layers, provider) so that each delete + * is not blocked by a reference from an object deleted later. On a clean boot there is nothing to remove and this + * is a handful of no-op calls, so "not found" is a success at every step. + * + * Coverage rests on every eBPF WFP object being tagged with EBPF_WFP_PROVIDER, which is what makes a purge by + * provider key exhaustive. Two properties keep that true: the provider GUID has never changed since it was + * introduced, and these filters are not persistent (no FWPM_FILTER_FLAG_PERSISTENT), so a reboot clears them and + * the only instance that can have left one behind is one from the current boot. Making these filters persistent + * would break this purge in two ways at once: it would need a sweep by callout GUID to catch objects tagged with + * an older provider GUID, and it would have to keep enumerating disabled filters, because BFE disables a + * persistent filter at boot when its provider declares no Windows service name, as this one does not. + * + * @param[in] engine_handle Open WFP engine handle to operate on. + * @retval STATUS_SUCCESS No eBPF WFP objects remain from a previous instance. + * @retval Other An object could not be removed, or the sweep could not be proven complete. The caller MUST fail + * initialization. + */ +_IRQL_requires_(PASSIVE_LEVEL) static NTSTATUS _net_ebpf_ext_purge_stale_wfp_objects(_In_ HANDLE engine_handle) +{ + NTSTATUS status = STATUS_SUCCESS; + const int max_retries = 10; + HANDLE filter_enum_handle = NULL; + HANDLE callout_enum_handle = NULL; + BOOLEAN is_in_transaction = FALSE; + BOOLEAN enumeration_drained = FALSE; + uint32_t stale_filter_count = 0; + size_t index; + + EBPF_EXT_LOG_ENTRY(); + + // Run in its own transaction so that aborting the initialization transaction cannot roll the purge back and + // leave the stale objects behind. Note that this holds for a purge that COMPLETES: the failure paths below + // deliberately abort this transaction, discarding whatever was reclaimed, because the caller refuses to start + // in that case and inert stale filters are harmless as long as no callout is registered. + status = FwpmTransactionBegin(engine_handle, 0); + EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmTransactionBegin", status); + is_in_transaction = TRUE; + + // Step 1: delete every filter tagged with the eBPF provider. + { + FWPM_FILTER_ENUM_TEMPLATE filter_enum_template = {0}; + filter_enum_template.providerKey = (GUID*)&EBPF_WFP_PROVIDER; + filter_enum_template.actionMask = 0xFFFFFFFF; // Ignore the filter's action type when enumerating. + + // Include disabled filters. BFE disables a filter at boot when its provider has no associated Windows + // service name, which this provider does not set, and a disabled filter is excluded by default. Such a + // filter cannot occur today because these filters are non-persistent, but if that ever changes a stale + // filter would otherwise be invisible here while still holding a dangling rawContext. + filter_enum_template.flags = FWP_FILTER_ENUM_FLAG_INCLUDE_DISABLED; + + status = FwpmFilterCreateEnumHandle(engine_handle, &filter_enum_template, &filter_enum_handle); + EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmFilterCreateEnumHandle", status); + + // A WFP enumerator is not live and always advances, so draining it terminates. The bound is defensive + // only: this runs during driver initialization, where spinning forever would be unrecoverable. + for (uint32_t batch = 0; batch < NET_EBPF_EXT_PURGE_MAX_ENUM_BATCHES; batch++) { + FWPM_FILTER** filters = NULL; + uint32_t filter_count = 0; + + status = FwpmFilterEnum( + engine_handle, filter_enum_handle, NET_EBPF_EXT_PURGE_ENUM_BATCH_SIZE, &filters, &filter_count); + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmFilterEnum", status); + goto Exit; + } + + if (filter_count == 0) { + if (filters != NULL) { + FwpmFreeMemory((void**)&filters); + } + enumeration_drained = TRUE; + break; + } + + for (uint32_t i = 0; i < filter_count; i++) { + uint64_t filter_id = filters[i]->filterId; + + for (int attempt = 1; attempt <= max_retries; attempt++) { + status = FwpmFilterDeleteById(engine_handle, filter_id); + if (NT_SUCCESS(status) || WFP_ERROR(status, FILTER_NOT_FOUND)) { + status = STATUS_SUCCESS; + break; + } + } + + if (!NT_SUCCESS(status)) { + // Count rather than bail: the log is far more useful with the full count of what could not + // be reclaimed, and the caller fails initialization either way. + stale_filter_count++; + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmFilterDeleteById", status); + } + } + + FwpmFreeMemory((void**)&filters); + } + + status = FwpmFilterDestroyEnumHandle(engine_handle, filter_enum_handle); + filter_enum_handle = NULL; + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmFilterDestroyEnumHandle", status); + goto Exit; + } + + if (!enumeration_drained) { + // The batch bound was exhausted, so the enumeration was never drained and an unknown number of eBPF + // filters was never examined. Any one of them may still carry a dangling rawContext, so this must fail + // rather than let the caller reach FwpsCalloutRegister. + status = STATUS_UNSUCCESSFUL; + EBPF_EXT_LOG_MESSAGE( + EBPF_EXT_TRACELOG_LEVEL_ERROR, + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, + "Enumeration of stale eBPF WFP filters did not complete. Refusing to initialize, because filters " + "that were never examined may still reference freed memory."); + goto Exit; + } + } + + if (stale_filter_count > 0) { + // The safety invariant: never proceed to register callouts while any eBPF filter with a dangling + // rawContext is still installed. + status = STATUS_UNSUCCESSFUL; + EBPF_EXT_LOG_MESSAGE_UINT32( + EBPF_EXT_TRACELOG_LEVEL_ERROR, + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, + "Stale eBPF WFP filters from a previous driver instance could not be deleted. Refusing to " + "initialize, because registering callouts while they are installed would use freed memory.", + stale_filter_count); + goto Exit; + } + + // Step 2: delete the callout management objects tagged with the eBPF provider. Nothing is registered at this + // point, so there is no corresponding FwpsCalloutUnregisterById to perform. + { + FWPM_CALLOUT_ENUM_TEMPLATE callout_enum_template = {0}; + callout_enum_template.providerKey = (GUID*)&EBPF_WFP_PROVIDER; + enumeration_drained = FALSE; + + status = FwpmCalloutCreateEnumHandle(engine_handle, &callout_enum_template, &callout_enum_handle); + EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmCalloutCreateEnumHandle", status); + + for (uint32_t batch = 0; batch < NET_EBPF_EXT_PURGE_MAX_ENUM_BATCHES; batch++) { + FWPM_CALLOUT** callouts = NULL; + uint32_t callout_count = 0; + + status = FwpmCalloutEnum( + engine_handle, callout_enum_handle, NET_EBPF_EXT_PURGE_ENUM_BATCH_SIZE, &callouts, &callout_count); + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmCalloutEnum", status); + goto Exit; + } + + if (callout_count == 0) { + if (callouts != NULL) { + FwpmFreeMemory((void**)&callouts); + } + enumeration_drained = TRUE; + break; + } + + for (uint32_t i = 0; i < callout_count; i++) { + GUID callout_key = callouts[i]->calloutKey; + + for (int attempt = 1; attempt <= max_retries; attempt++) { + status = FwpmCalloutDeleteByKey(engine_handle, &callout_key); + if (NT_SUCCESS(status) || WFP_ERROR(status, CALLOUT_NOT_FOUND)) { + status = STATUS_SUCCESS; + break; + } + } + + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmCalloutDeleteByKey", status); + FwpmFreeMemory((void**)&callouts); + goto Exit; + } + } + + FwpmFreeMemory((void**)&callouts); + } + + status = FwpmCalloutDestroyEnumHandle(engine_handle, callout_enum_handle); + callout_enum_handle = NULL; + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmCalloutDestroyEnumHandle", status); + goto Exit; + } + + if (!enumeration_drained) { + // As with the filters above: an incomplete sweep cannot be reported as a successful purge. A stale + // callout object left behind would also block the FwpmCalloutAdd calls that follow. + status = STATUS_UNSUCCESSFUL; + EBPF_EXT_LOG_MESSAGE( + EBPF_EXT_TRACELOG_LEVEL_ERROR, + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, + "Enumeration of stale eBPF WFP callouts did not complete. Refusing to initialize."); + goto Exit; + } + } + + // Step 3: delete the sub-layers this driver version knows about. There is no sub-layer enumeration by + // provider here because the static table is the authoritative list for this version; a sub-layer created by + // a different version would be removed by that version's own purge. + for (index = 0; index < EBPF_COUNT_OF(_net_ebpf_ext_sublayers); index++) { + for (int attempt = 1; attempt <= max_retries; attempt++) { + status = FwpmSubLayerDeleteByKey(engine_handle, _net_ebpf_ext_sublayers[index].sublayer_guid); + if (NT_SUCCESS(status) || WFP_ERROR(status, SUBLAYER_NOT_FOUND)) { + status = STATUS_SUCCESS; + break; + } + } + + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmSubLayerDeleteByKey", status); + goto Exit; + } + } + + // Step 4: delete the provider itself, now that nothing references it. + for (int attempt = 1; attempt <= max_retries; attempt++) { + status = FwpmProviderDeleteByKey(engine_handle, &EBPF_WFP_PROVIDER); + if (NT_SUCCESS(status) || WFP_ERROR(status, PROVIDER_NOT_FOUND)) { + status = STATUS_SUCCESS; + break; + } + } + + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmProviderDeleteByKey", status); + goto Exit; + } + + status = FwpmTransactionCommit(engine_handle); + EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmTransactionCommit", status); + is_in_transaction = FALSE; + +Exit: + + // Enumeration handles are closed on the success path above; these fire only when an error left one open. + if (filter_enum_handle != NULL) { + NTSTATUS destroy_status = FwpmFilterDestroyEnumHandle(engine_handle, filter_enum_handle); + if (!NT_SUCCESS(destroy_status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmFilterDestroyEnumHandle", destroy_status); + } + } + + if (callout_enum_handle != NULL) { + NTSTATUS destroy_status = FwpmCalloutDestroyEnumHandle(engine_handle, callout_enum_handle); + if (!NT_SUCCESS(destroy_status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmCalloutDestroyEnumHandle", destroy_status); + } + } + + if (is_in_transaction) { + NTSTATUS abort_status = FwpmTransactionAbort(engine_handle); + if (!NT_SUCCESS(abort_status)) { + EBPF_EXT_LOG_NTSTATUS_API_FAILURE( + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmTransactionAbort", abort_status); + } + } + + EBPF_EXT_RETURN_NTSTATUS(status); +} + NTSTATUS net_ebpf_extension_initialize_wfp_components(_Inout_ void* device_object) /* ++ @@ -822,6 +1124,19 @@ net_ebpf_extension_initialize_wfp_components(_Inout_ void* device_object) EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmEngineOpen", status); is_engine_opened = TRUE; + // Remove any WFP objects left behind by a previous driver instance before adding our own. This MUST happen + // before the first FwpsCalloutRegister below: a stale filter still holds the address of a filter context that + // the previous instance freed, so registering a callout it references would hand WFP a dangling pointer. + // Failure here is fatal by design -- see _net_ebpf_ext_purge_stale_wfp_objects. + status = _net_ebpf_ext_purge_stale_wfp_objects(_fwp_engine_handle); + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_MESSAGE( + EBPF_EXT_TRACELOG_LEVEL_ERROR, + EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, + "Failed to purge stale eBPF WFP objects from a previous driver instance."); + goto Exit; + } + status = FwpmTransactionBegin(_fwp_engine_handle, 0); EBPF_EXT_BAIL_ON_API_FAILURE_STATUS(EBPF_EXT_TRACELOG_KEYWORD_EXTENSION, "FwpmTransactionBegin", status); is_in_transaction = TRUE; diff --git a/netebpfext/sys/net_ebpf_ext_drv.c b/netebpfext/sys/net_ebpf_ext_drv.c index c9ef580f18..38f83148d8 100644 --- a/netebpfext/sys/net_ebpf_ext_drv.c +++ b/netebpfext/sys/net_ebpf_ext_drv.c @@ -51,8 +51,6 @@ _net_ebpf_ext_driver_uninitialize_objects() net_ebpf_ext_uninitialize_ndis_handles(); - ebpf_ext_trace_terminate(); - if (_net_ebpf_ext_device != NULL) { WdfObjectDelete(_net_ebpf_ext_device); } @@ -63,6 +61,10 @@ static _Function_class_(EVT_WDF_DRIVER_UNLOAD) _IRQL_requires_same_ { UNREFERENCED_PARAMETER(driver_object); _net_ebpf_ext_driver_uninitialize_objects(); + + // Tracing is torn down last, so that everything above can still be traced. It is started by DriverEntry, and + // every path that leaves the driver loaded-and-running ends here. + ebpf_ext_trace_terminate(); } // @@ -128,8 +130,24 @@ _net_ebpf_ext_driver_initialize_objects(_Inout_ DRIVER_OBJECT* driver_object, _I goto Exit; } - // TODO: https://github.com/microsoft/ebpf-for-windows/issues/521 - (void)net_ebpf_extension_initialize_wfp_components(_net_ebpf_ext_driver_device_object); + // A failure here means no eBPF network hook can work, so fail the driver start rather than running in a + // silently inert state. DriverEntry calls _net_ebpf_ext_driver_uninitialize_objects() and returns this + // status, so the service start fails with a real error code. + // + // TODO: https://github.com/microsoft/ebpf-for-windows/issues/521 - this driver still does not register for + // BFE service status before adding WFP objects. Failing the start makes that gap visible instead of hiding + // it, but it also means a start attempted while BFE is unavailable now fails outright; registering for BFE + // status notifications and adding the WFP objects once BFE is ready is the complete fix. + status = net_ebpf_extension_initialize_wfp_components(_net_ebpf_ext_driver_device_object); + if (!NT_SUCCESS(status)) { + EBPF_EXT_LOG_MESSAGE_NTSTATUS( + EBPF_EXT_TRACELOG_LEVEL_ERROR, + EBPF_EXT_TRACELOG_KEYWORD_BASE, + "net_ebpf_extension_initialize_wfp_components failed. The driver cannot provide any eBPF network " + "hooks, so the start is failed rather than left silently non-functional.", + status); + goto Exit; + } status = net_ebpf_ext_register_providers(); if (!NT_SUCCESS(status)) { @@ -175,5 +193,12 @@ DriverEntry(_In_ DRIVER_OBJECT* driver_object, _In_ UNICODE_STRING* registry_pat } EBPF_EXT_LOG_EXIT(); + + if (!NT_SUCCESS(status)) { + // Terminate tracing only after the exit log above, so that a failed start is fully traced. This is the + // reason _net_ebpf_ext_driver_uninitialize_objects() does not terminate tracing itself. + ebpf_ext_trace_terminate(); + } + return status; } From 37875df478e6d98cc1619ebf609c12839b883c14 Mon Sep 17 00:00:00 2001 From: Michael Agun Date: Wed, 12 Aug 2026 05:12:04 -0700 Subject: [PATCH 2/3] Bump usersim for WFP mock enumeration APIs The netebpfext stale-WFP-object purge discovers objects by enumerating them, which the user-mode WFP mock could not do, and it relies on the mock modelling provider identity and object references to be testable at all. TEMPORARY: this also points the external/usersim submodule URL at a fork so that CI can build this pull request before the usersim change merges. Both the URL and the commit must be moved back to microsoft/usersim, at the merged commit, before this merges. Depends on microsoft/usersim#325. --- .gitmodules | 2 +- external/usersim | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/.gitmodules b/.gitmodules index 1354164cdd..49d41181a8 100644 --- a/.gitmodules +++ b/.gitmodules @@ -20,7 +20,7 @@ fetchRecurseSubmodules = no [submodule "external/usersim"] path = external/usersim - url = https://github.com/microsoft/usersim.git + url = https://github.com/mikeagun/usersim.git [submodule "external/ebpf-extension-common"] path = external/ebpf-extension-common url = https://github.com/microsoft/ebpf-extension-common.git diff --git a/external/usersim b/external/usersim index dcff48ce18..6213ba2a20 160000 --- a/external/usersim +++ b/external/usersim @@ -1 +1 @@ -Subproject commit dcff48ce18503f996848436cdba14e0299b0b262 +Subproject commit 6213ba2a20851882d3c1f1c6ccd2650edfb7a010 From 9205b09f636fa97e635aec0c81481eebfa4d80de Mon Sep 17 00:00:00 2001 From: Michael Agun Date: Thu, 13 Aug 2026 23:01:54 -0700 Subject: [PATCH 3/3] netebpfext: test the stale WFP object purge Adds the positive case, where an unload that leaves filters installed is followed by a start that purges them and succeeds, and the negative case that guards the safety invariant: a stale filter that cannot be deleted must make initialization fail rather than register callouts it still references. The helper gains a wfp_is_initialized() accessor because its constructor cannot use REQUIRE. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5196c1b8-883f-440e-8846-e5fd79db2f5d --- tests/netebpfext_unit/netebpf_ext_helper.h | 8 ++ tests/netebpfext_unit/netebpfext_unit.cpp | 92 ++++++++++++++++++++++ 2 files changed, 100 insertions(+) diff --git a/tests/netebpfext_unit/netebpf_ext_helper.h b/tests/netebpfext_unit/netebpf_ext_helper.h index f4e4465cf9..862c5f2370 100644 --- a/tests/netebpfext_unit/netebpf_ext_helper.h +++ b/tests/netebpfext_unit/netebpf_ext_helper.h @@ -132,6 +132,14 @@ typedef class _netebpf_ext_helper } } + // Whether net_ebpf_extension_initialize_wfp_components() succeeded. The constructor cannot use REQUIRE, so + // tests that exercise initialization failure check this instead. + bool + wfp_is_initialized() const + { + return wfp_initialized; + } + private: bool trace_initiated = false; bool ndis_handle_initialized = false; diff --git a/tests/netebpfext_unit/netebpfext_unit.cpp b/tests/netebpfext_unit/netebpfext_unit.cpp index d9947d5130..042b556917 100644 --- a/tests/netebpfext_unit/netebpfext_unit.cpp +++ b/tests/netebpfext_unit/netebpfext_unit.cpp @@ -1015,7 +1015,99 @@ TEST_CASE("wfp_filter_delete_failure_unload_deletes_stale_filter", "[netebpfext] REQUIRE(usersim_fwp_get_fwpm_filter_count() == 0); } +// Verifies that a start following an unload that left filters installed purges them: the stale filters pin the +// provider, sub-layers and callout objects, so without the purge FwpmProviderAdd fails with FWP_E_ALREADY_EXISTS +// and the driver comes up with no WFP components at all. +TEST_CASE("wfp_stale_objects_purged_on_initialize", "[netebpfext][wfp_cleanup]") +{ + if (cxplat_fault_injection_is_enabled()) { + return; + } + + ebpf_extension_data_t npi_specific_characteristics = { + .header = EBPF_ATTACH_CLIENT_DATA_HEADER_VERSION, + }; + test_sock_addr_client_context_header_t client_context_header = {0}; + client_context_header.context.base.desired_attach_types = {BPF_CGROUP_INET4_CONNECT, BPF_CGROUP_INET6_CONNECT}; + test_sock_addr_client_context_t* client_context = &client_context_header.context; + + { + netebpf_ext_helper_t helper( + &npi_specific_characteristics, + (_ebpf_extension_dispatch_function)netebpfext_unit_invoke_sock_addr_program, + (netebpfext_helper_base_client_context_t*)client_context); + + REQUIRE(usersim_fwp_get_fwpm_filter_count() > 0); + + // Fail every delete so the filters survive detach and unload. + usersim_fwp_set_filter_delete_failure_count(UINT32_MAX); + helper.detach_hook_client(); + + // End of scope runs unload, which cannot delete the filters or the objects they reference. + } + + REQUIRE(usersim_fwp_get_fwpm_filter_count() > 0); + + // Deletes succeed again, which is what a real restart looks like once whatever blocked them has cleared. + usersim_fwp_set_filter_delete_failure_count(0); + + { + // Starting again must purge the stale objects and initialize successfully. + netebpf_ext_helper_t helper; + + REQUIRE(helper.wfp_is_initialized()); + REQUIRE(usersim_fwp_get_fwpm_filter_count() == 0); + } +} + +// Verifies the safety invariant behind the purge: if a stale filter cannot be deleted, initialization must fail. +// Such a filter holds the address of a filter context the previous driver instance freed, so registering a callout +// it references would hand WFP a dangling pointer. A partial purge followed by a successful start is strictly worse +// than refusing to start. +TEST_CASE("wfp_initialize_fails_when_stale_filter_cannot_be_purged", "[netebpfext][wfp_cleanup]") +{ + if (cxplat_fault_injection_is_enabled()) { + return; + } + + ebpf_extension_data_t npi_specific_characteristics = { + .header = EBPF_ATTACH_CLIENT_DATA_HEADER_VERSION, + }; + test_sock_addr_client_context_header_t client_context_header = {0}; + client_context_header.context.base.desired_attach_types = {BPF_CGROUP_INET4_CONNECT, BPF_CGROUP_INET6_CONNECT}; + test_sock_addr_client_context_t* client_context = &client_context_header.context; + + { + netebpf_ext_helper_t helper( + &npi_specific_characteristics, + (_ebpf_extension_dispatch_function)netebpfext_unit_invoke_sock_addr_program, + (netebpfext_helper_base_client_context_t*)client_context); + + REQUIRE(usersim_fwp_get_fwpm_filter_count() > 0); + + usersim_fwp_set_filter_delete_failure_count(UINT32_MAX); + helper.detach_hook_client(); + } + + REQUIRE(usersim_fwp_get_fwpm_filter_count() > 0); + + { + // Deletes still fail, so the purge cannot remove the stale filters and the start must be refused. + netebpf_ext_helper_t helper; + + REQUIRE_FALSE(helper.wfp_is_initialized()); + } + + // The stale filters are still installed, as nothing was able to delete them. + REQUIRE(usersim_fwp_get_fwpm_filter_count() > 0); + + // Clear them directly; the next start's purge reclaims the objects they were pinning. + usersim_fwp_set_filter_delete_failure_count(0); + usersim_fwp_clear_fwpm_filters(); +} + #pragma endregion cgroup_sock_addr + #pragma region sock_ops typedef enum _sock_ops_test_action