Skip to content

[meshnet] multiplex tcp streams for a node pair - #733

Open
kraney wants to merge 7 commits into
openconfig:mainfrom
kraney:meshnet-multiplex
Open

[meshnet] multiplex tcp streams for a node pair#733
kraney wants to merge 7 commits into
openconfig:mainfrom
kraney:meshnet-multiplex

Conversation

@kraney

@kraney kraney commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Note: this PR stacks on top of #732 , please merge that one first. The incremental changes start at sha ceb9511

  1. Created nodeStreamManager (stream_manager.go):
      • Keyed by nodeStreamKey{ topoNs string, peerIP string }.
      • Opens only one gRPC connection (grpc.Dial) and only one SendToStream 
         bidirectional streaming RPC per (peerIP, topoNs) pair.
      • Includes reference counting (GetOrCreateStream / ReleaseStream). When the last 
         wire for a topology targeting a peer node is removed, the shared stream and gRPC 
         connection close gracefully.
      • Features a high-capacity buffered channel (10,000 packet queue) and dedicated sender 
         worker loop with automatic reconnect logic.
  2. Updated TAP Reader Threads (grpcwire.go):
      • Modified RecvFrmLocalPodThread so TAP readers no longer execute individual grpc.Dial 
         or SendToStream calls.
      • Each TAP reader now acquires the shared topology stream via 
         nodeStream := streamMgr.GetOrCreateStream(wire.TopoNamespace, wire.PeerNodeIP) 
        and multiplexes packets into nodeStream.Send(payload).
  3. Per-Topology Isolation:
      • Topology topo-A and topology topo-B running between the same pair of physical 
          nodes maintain completely separate gRPC connections and streams.

Also:
The base design had a lot of single-interface RPCs
to update k8s and remote meshnets. For large numbers
of links it costs a lot of serial round trip wait time.
This batches things and pipelines them so that we avoid
a lot of unnecessary overhead delay at start time.

kraney added 7 commits July 31, 2026 20:31
Replace the veth pair and libpcap packet capture implementation in meshnetd
with persistent Linux TAP devices created directly inside container network
namespaces.

Reason for Change:
- Eliminates packet loss under high traffic loads caused by libpcap's kernel-
  to-userspace drop policy.
- Provides native Linux socket buffer flow control and backpressure (`txqueuelen`)
  up to the container application when gRPC processing falls behind.
- Guarantees protocol transparent packet handling (including LACP and LLDP frames).
- Ensures process crash resilience: setting `TUNSETPERSIST=1` keeps TAP interfaces
  alive in the container netns across daemon restarts without link flaps.
- Synchronizes CNI plugin pod readiness with complete end-to-end gRPC wire setup,
  preventing test/ping race conditions upon pod startup.
- Eliminates CGO and `libpcap-dev` build dependencies, producing a pure Go static
  binary (`CGO_ENABLED=0`).

Key Changes:
- Added `CreateOrAttachTAP` in `wireutil` using `TUNSETIFF` & `TUNSETPERSIST`.
- Replaced `pcap.Handle` with `*os.File` in `gwire_map.go`, `grpcwire.go`, and `handler.go`.
- Updated `ReconcilePodLinks` in `controller.go` to use TAP interfaces directly without
  host-side veth creation.
- Updated `GRPCWireExists` and CNI `cmdAdd` readiness check to block until gRPC wire
  handshakes are fully established.
- Removed `libpcap-dev` and updated Dockerfile to build `meshnetd` with `CGO_ENABLED=0`.
Add package & public method comments, and format Go
  1. Bidirectional Streaming (SendToStream) Receiver (handler.go):
      • Implemented SendToStream(stream mpb.WireProtocol_SendToStreamServer) error.
      • In a loop, stream.Recv() continuously ingests incoming mpb.Packet frames and writes them directly to the
      destination TAP interface (wrHandle.Write(pkt.Frame)), bypassing per-packet unary RPC overhead.
  2. Streaming Sender & Auto-Reconnect (grpcwire.go):
      • Updated RecvFrmLocalPodThread to establish a persistent client stream (wireClient.SendToStream(ctx)).
      • Frames read from the local TAP interface are streamed out via st.Send(payload) without blocking for individual
      RPC responses.
      • If the stream encounters a network or peer reset, RecvFrmLocalPodThread automatically clears its stream handle
      and transparently re-establishes SendToStream on the next packet.
  3. HTTP/2 Window Size & Buffer Tuning:
      • Server Configuration (meshnet.go):
          • Stream window size: 4 MB (grpc.InitialWindowSize(4 * 1024 * 1024)).
          • Connection window size: 16 MB (grpc.InitialConnWindowSize(16 * 1024 * 1024)).
          • Max message payload: 64 MB (grpc.MaxRecvMsgSize / grpc.MaxSendMsgSize).
      • Client Configuration (grpcwire.go):
          • Configured matching initial window size and message payload limits on client grpc.Dial.
  1. Created nodeStreamManager (stream_manager.go):
      • Keyed by nodeStreamKey{ topoNs string, peerIP string }.
      • Opens only one gRPC connection (grpc.Dial) and only one SendToStream bidirectional streaming RPC per (peerIP,
      topoNs) pair.
      • Includes reference counting (GetOrCreateStream / ReleaseStream). When the last wire for a topology targeting a
      peer node is removed, the shared stream and gRPC connection close gracefully.
      • Features a high-capacity buffered channel (10,000 packet queue) and dedicated sender worker loop with automatic
      reconnect logic.
  2. Updated TAP Reader Threads (grpcwire.go):
      • Modified RecvFrmLocalPodThread so TAP readers no longer execute individual grpc.Dial or SendToStream calls.
      • Each TAP reader now acquires the shared topology stream via nodeStream := streamMgr.GetOrCreateStream(wire.
      TopoNamespace, wire.PeerNodeIP) and multiplexes packets into nodeStream.Send(payload).
  3. Per-Topology Isolation:
      • Topology topo-A and topology topo-B running between the same pair of physical nodes maintain completely
      separate gRPC connections and streams.
The base design had a lot of single-interface RPCs
to update k8s and remote meshnets. For large numbers
of links it costs a lot of serial round trip wait time.
This batches things and pipelines them so that we avoid
a lot of unnecessary overhead delay at start time.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant