Rich Gibbs

Proving Inter-VLAN Routing, After Three Failures That Weren't Routing

homelab · networking · vlan · sg300 · networkmanager · dnsmasq · dns

In session 25, testing inter-VLAN routing from a real client turned up three failures that each looked like routing and weren't: the SG300 firmware image had no DHCP server, NetworkManager wiped an address added with ip addr add, and dnsmasq's localservice setting ignored queries from a VLAN reached by static route. A TTL drop from 64 to 63 proved the switch was routing. DNS was fixed narrowly by pointing VLAN 20 at a public resolver instead of disabling localservice. ACLs were still not done.

Session 24 built inter-VLAN routing on the SG300 switch but never tested it from an actual client. Session 25 was the test. It failed three times, and every failure looked like a routing problem. None of them were. The routing had been correct the whole time.

Writing this down so the next time something "doesn't route," I check the boring layers first.

The switch can't hand out addresses, and it never could

Session 24 ended with the DHCP pool command coming back as unrecognized. On firmware from 2011, that read like a syntax problem.

Rather than trying variations, I asked the command tree what actually exists. ip dhcp ? listed three things: information, relay and tftp-server. No pool. No server. This firmware image simply has no DHCP server feature. It was never a syntax problem. My note at the time says DHCP server support arrived in later SG300 firmware, but what I actually saw was this one image, and that's the only thing I'm sure of.

The options:

  • DHCP relay. The edge router would then need pools for subnets it has no interface in, which means more config on top of firmware that had already let me down once.
  • Static addressing. Zero new config, unblocks everything now.
  • A DHCP server on the Proxmox node. That's the proper long-term answer, but it's chicken-and-egg: Proxmox has to move to VLAN 10 first.

I went static. The reasoning is worth keeping: DHCP was never the goal. The goal is proving routing and then writing ACL policy on top of it. Static addressing tests routing just as well and takes a failure mode off the table. Proxmox is statically addressed anyway, so its eventual move to VLAN 10 doesn't depend on DHCP either.

Lesson: check whether a feature exists before debugging its syntax. <command> ? answers that in one line.

NetworkManager quietly undid my work

On the test laptop (Linux Mint) I set up VLAN 20 the obvious way. I flushed the wired interface, added an address with ip addr add, and added a default route with ip route add. ip route showed both.

Then every ping failed with "Network is unreachable", including the ping to the gateway on the same subnet.

That last part is the tell. Pinging a directly connected gateway needs no default route, no switch routing and no edge router. It needs ARP. If that fails with "Network is unreachable", the route table isn't what you think it is.

ip -br address showed only an IPv6 link-local address on the interface. The IPv4 address was gone. NetworkManager had reasserted control of the interface and wiped the manual config seconds after I applied it.

The fix was to configure through NetworkManager instead of behind its back: nmcli con mod on the existing wired connection profile, setting the method to manual plus the address, gateway and DNS, then nmcli con up. The shape of it, with placeholders rather than the real values:

# Example only, placeholders in angle brackets
sudo nmcli con mod "<connection>" ipv4.method manual \
  ipv4.addresses <addr/prefix> ipv4.gateway <gateway> ipv4.dns <dns>
sudo nmcli con up "<connection>"

Lesson: on a NetworkManager system, ip addr add is a temporary override, not a configuration. Better to learn that on a test laptop than on something that matters.

The TTL is the proof

Once the address held, with WiFi off and the cable in a VLAN 20 port, I ran a ladder of four pings: the VLAN 20 gateway on the switch, the switch's management address, the edge router, then an internet host. All four came back with 0% loss.

The TTLs were what actually told me something: 64, 64, 63, 57.

The first two stay at 64 because both the VLAN 20 SVI and the VLAN 1 SVI are interfaces on the switch itself. The packet is delivered locally, not forwarded. The drop to 63 at the edge router is the switch decrementing TTL as it routes between subnets. That single digit is the difference between "these hosts are on one flat network" and "the switch routed my traffic between two subnets."

The path: laptop on VLAN 20, to the switch SVI, the switch routes VLAN 20 to VLAN 1, then the edge router, then NAT, upstream WiFi and the internet.

The return routes added to the router in session 24 are why the last hop worked. My reasoning is that without them, the first three pings would pass and the internet one would fail. They were already in place, so I never saw that asymmetric-routing failure. It's a prediction, not something I watched happen.

Read the TTLs, not just the successes.

Ping works, DNS doesn't

Pinging a public hostname failed with "Name or service not known", while pinging an internet IP worked. So routing was fine and only name resolution was broken, a clean, isolated layer.

resolvectl status showed the laptop correctly pointed at the edge router for DNS. Querying the router directly with nslookup timed out. The same host that answers ICMP wasn't answering on port 53, which rules routing out completely: the packets were arriving.

On the router, uci show had dnsmasq set to localservice='1', and logread showed dnsmasq repeatedly logging "Ignoring query from non-local network". It told me exactly what it was doing. With localservice=1, dnsmasq only answers clients on directly attached subnets. VLAN 20 is reached by a static route, not by an interface, so its queries were classed as non-local and dropped. The firewall was innocent; the lan zone input policy was ACCEPT.

Two fixes:

  • Blunt. Set localservice='0'. It works instantly, but it doesn't scope anything. It removes the check entirely and leaves only the WAN firewall between this resolver and the internet. Open resolvers get scanned for and abused in DNS amplification attacks. That's one layer instead of two, to solve a local problem.
  • Narrow. Point VLAN 20 clients at a public resolver instead. Zero change to the router's security posture. It costs local name resolution on that VLAN, which for an isolated lab segment is arguably correct. A quarantined segment having no view of the internal namespace is a feature.

I tried both, then set localservice back to 1. DNS still resolved, which proves the client-side change is what's carrying it.

There's a trap I want future me to remember. It's tempting to add a secondary IP in the VLAN 20 subnet on the router's LAN bridge so dnsmasq treats the subnet as local and localservice=1 can stay. Don't. The router would then consider that subnet directly connected, the connected route would beat the static route via the switch, and the return path would silently break. Elegant-looking, quietly destructive. I didn't try it. That's reasoning, and it's enough.

Lesson: don't disable a security control globally to solve a local problem. Ask what the control is for, then make the narrowest change that meets the requirement.

Longer term, DNS is a service and belongs on a server, not on the edge router. A resolver on the Proxmox node would give per-VLAN policy, local names and, for the SOC track, DNS query logging, which is one of the highest-value telemetry sources there is. That's the plan, not something that exists yet.

Cleanup, and the trap that keeps coming back

I took the test scaffolding out: WiFi radio back on, the wired connection back to automatic addressing, the cable moved back to a VLAN 1 port. I kept the router's static routes for the two new VLAN subnets, and localservice='1', the secure default.

After cleanup, both wired and WiFi held default routes, with wired winning on metric (100 vs 600). Switch management was still reachable.

That two-path state is the same trap from session 23 and from Part 2 of session 24. Turn WiFi off before any isolation test, or the result lies.

Where it stood

Proven: inter-VLAN routing from VLAN 20 through VLAN 1 to the router and out to the internet, verified by TTL; DNS from VLAN 20 via a public resolver; the router's return routes in its kernel table.

Known limits: no DHCP server on this switch firmware, so static addressing on VLANs 10 and 20, and no local name resolution on the routed VLANs, by choice.

Not done:

  • ACLs. VLAN 20 could reach VLAN 10, the switch management interface, the router, and every VLAN 1 host, including the Proxmox node and both VMs. That is routing, not isolation. Isolation has to be written as explicit policy, and that's the next session.
  • The Proxmox node was still on VLAN 1.
  • No reboot durability test.
  • NTP on the SG300 still unconfigured.

The thing to remember from session 25 is that three "routing problems" were really a missing feature, a network manager, and a resolver doing exactly what it was told. And the routing working is only the halfway point.