Skip to content

topology.tuple.deserialization.strict.enable does not terminate the worker for failures wrapping an IOException #9157

Description

@L1nq0

Observed with topology.tuple.deserialization.strict.enable set to true (the flag from #9094):

When an incoming tuple fails to deserialize and the failure carries an IOException in its cause chain, the receiving worker does not terminate. The connection is closed, the batch being read, including tuples that decoded fine, is discarded, and the peer reconnects. Failures without an IOException in the chain, such as unknown task or stream ids, do terminate the worker.

The path: DeserializingConnectionCallback.recv rethrows the failure in strict mode, the exception reaches StormServerHandler.exceptionCaught, and Utils.handleUncaughtException is called with an allowed-exception set containing only IOException. The check walks the cause chain, and both ZstdUtils.decompress and KryoTupleDeserializer.deserializeTuple wrap IOException in a RuntimeException, so compressed-frame failures (a broken frame, or decompression over topology.tuple.compression.max.decompressed.bytes) match the exemption. The handler logs, closes the channel, and returns.

Two consequences beyond the worker staying alive:

  • The whole batch in flight is lost, including tuples that decoded fine.
  • deserializationFailures is not incremented, because the strict path rethrows before the counter.

Two directions:

Follow-up from the review discussion in #9094. Related: #9077.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions