Skip to content

bug: Podman driver picks the wrong host IP for the supervisor callback on multi-homed Linux hosts (regression from #2942) #3412

Description

@jgarciao

User Story

As a user running the local Podman driver on a multi-homed Linux laptop (wired dock + Wi-Fi + Tailscale), I want openshell sandbox create to work without hand-editing the gateway config or bringing interfaces down.

Problem Statement

On a host with multiple network interfaces, the Podman driver and Podman disagree on which host IP host.containers.internal points to, so the in-container supervisor can't reach the gateway and provisioning fails.

  • The gateway binds its callback listener to the default-route interface plus loopback: 192.168.1.27:17670 and 127.0.0.1:17670.
  • The driver tells the supervisor to dial host.containers.internal:17670 and injects --add-host host.containers.internal:host-gateway.
  • Podman resolves host-gateway to a different interface — here the Tailscale address 100.64.0.2 (wt0, CGNAT 100.64.0.0/10), where nothing listens. Bringing Wi-Fi down didn't help; Podman then used the Tailscale IP instead of the wired one.
  • The supervisor's policy fetch fails, it exits 1, the sandbox container follows, and the sandbox enters Error.

The reported error is also misleading: ContainerExited: Container exited with code 0 refers to the sandbox container reacting to its supervisor dying. The real failure is only in the openshell-supervisor-<id> container logs.

Impact / Why This Matters

  • Consequence: sandbox creation fails outright on multi-homed hosts (dock + Wi-Fi + VPN/mesh laptops are common), and the surfaced error points at the wrong container, so it's hard to diagnose.

  • Workaround: set host_gateway_ip = "127.0.0.1" under [openshell.drivers.podman] in ~/.config/openshell/gateway.toml, then systemctl --user restart openshell-gateway. This works because the supervisor uses --network host, so loopback reaches the gateway's 127.0.0.1:17670 listener, and both host.containers.internal and 127.0.0.1 are cert SANs. The sandbox then reaches Ready.

    [openshell.drivers.podman]
    host_gateway_ip = "127.0.0.1"
  • Why insufficient: it relies on an undocumented field and non-obvious reasoning about host-network loopback; it isn't discoverable from the error. The default should just work on multi-homed hosts.

Confirmed Regression (#2942)

Confirmed by version bisection: this is a regression introduced by PR #2942 (RFC 0012 sandbox architecture).

  • 0.0.116 (does not contain feat(isolation): implement the RFC 0012 sandbox architecture #2942; tagged 2026-08-28): openshell sandbox create reaches a working sandbox shell with default config (no host_gateway_ip) on this multi-homed host. ✅
  • 0.0.117-dev / current main (contains feat(isolation): implement the RFC 0012 sandbox architecture #2942, merged 2026-09-16 as c1f2e7189, first released in v0.1.0-pre.2): the same default config fails with Error, and only the host_gateway_ip = "127.0.0.1" workaround makes it succeed. ❌
  • The gateway binds the same two listeners (192.168.1.27 + 127.0.0.1) on both versions, so the DefaultRouteInterface listener is not the trigger — the change is the callback path.

Mechanism:

Acceptance Criteria

  • A multi-homed Linux host reaches Ready with the Podman driver using default config (no manual host_gateway_ip).
  • The supervisor reaches the gateway regardless of which interface Podman picks for host-gateway.
  • When the supervisor can't reach the gateway, the phase error reports its failure reason, not ContainerExited: code 0.
  • Covered by a test, or the multi-homed host_gateway_ip guidance is documented.

Reproduction Steps

  1. On a Linux host whose interfaces differ from Podman's host-gateway resolution (e.g. wired dock + Wi-Fi + Tailscale), configure the Podman driver.
  2. Run openshell sandbox create --provider <any>.
  3. The sandbox enters Error during provisioning.

Confirm the mismatch:

ss -ltn | grep 17670
#   LISTEN ... 192.168.1.27:17670   (default-route iface)
#   LISTEN ...    127.0.0.1:17670

podman run --rm --network host --add-host host.containers.internal:host-gateway \
  registry.fedoraproject.org/fedora-minimal:latest getent hosts host.containers.internal
#   100.64.0.2  host.containers.internal   <-- Tailscale wt0, nothing listening there

Environment

  • OpenShell: 0.0.117-dev.167+g7e7a8d561 (development build; regression absent in 0.0.116), installed with:
    curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | OPENSHELL_VERSION=dev sh
  • OS / kernel: Fedora Linux 44 (Workstation) / 7.2.5-200.fc44.x86_64
  • Podman: 5.8.4 (netavark)
  • Driver: podman (rootless); gateway runs as the openshell-gateway systemd user service (native install, not containerized)
  • Network: multi-homed — wired dock 192.168.1.27, Wi-Fi 192.168.1.75 (same subnet), Tailscale wt0 in 100.64.0.0/10

Related

Related: #2540, #1952, and the macOS cases #1519 / #1634. Distinct from #1909 (containerized gateway on Fedora 44).

Logs

Supervisor container (openshell-supervisor-<id>) — the actual failure:

WARN openshell_supervisor: Policy fetch failed, retrying
Error: × Policy fetch failed after 5 attempts: failed to connect to OpenShell server

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions