Posts

Showing posts with the label High Availability

FHRP Explained: HSRP, VRRP and GLBP for Gateway Redundancy

A host knows exactly one default gateway. If that router fails, the host does not look for another — it simply cannot reach anything off its own subnet. First Hop Redundancy Protocols solve this by letting two or more routers share a single virtual gateway address, so the host's configuration never has to change and failover happens without the host knowing. The problem Redundant links, redundant switches, and redundant uplinks all count for nothing if every host on the segment points at one router's address. That address is a single point of failure sitting in the configuration of every device. Changing it on failure is not an option: hosts learn it from DHCP and hold it for the lease duration, and statically configured devices never change it at all. The gateway address has to stay constant while the router behind it changes. How the virtual gateway works Two or more routers share a virtual IP address and a virtual MAC address. Hosts are configured — usually via DH...

Fail-Open vs Fail-Closed: Choosing Which Way a Control Breaks

Every control fails eventually, and it fails in one of two directions. Fail-open means access is permitted when the control stops working — availability is preserved, security is lost. Fail-closed , or fail-secure, means access is denied — security is preserved, availability is lost. Choosing is a business decision, and defaults frequently choose for you. The general principle and its exception Security controls should generally fail closed. A control that stops enforcing when it breaks provides an attacker with an obvious strategy: break the control. An authentication system that admits everyone when its database is unreachable is not an authentication system. The exception is life safety , and it overrides everything. Doors on emergency exit routes must unlock when power or control fails, because people trapped in a burning building is a worse outcome than an unauthorized entry. Fire codes require this and they are not negotiable — a security design that conflic...

Active/Active vs Active/Passive: Site Resilience and Recovery Models

Redundancy comes in two shapes, and the choice between them determines cost, complexity and how long recovery takes. Active/active runs all nodes serving traffic simultaneously. Active/passive runs one while the other stands ready. Everything else in resilience design follows from that distinction. Active/passive One node handles everything; the standby monitors and takes over on failure. Failover may be automatic through a health check and a virtual address, or manual. The advantages are simplicity and predictability. There is one authoritative node, so no question of conflicting writes, and capacity planning is straightforward because a single node carries the full load by definition. The costs are real. The standby is paid for and idle, failover takes time — seconds to minutes, occasionally longer — and sessions in flight are usually lost. Worst of all, an untested standby frequently does not work: configuration drifts, a licence expires, a certificate lapses, and ...

Load Balancing and Session Persistence: Algorithms, Affinity and Health Checks

A load balancer distributes requests across several backend servers, which adds capacity and removes any one server as a single point of failure. The complication is state: if a user's session lives in memory on one server, sending their next request elsewhere logs them out. Session persistence is the workaround, and designing so you do not need it is the better answer. The algorithms Round robin sends each request to the next server in turn. Simple, and it assumes every server and every request are equivalent — which they are not. Weighted round robin assigns proportions, so a server with twice the capacity receives twice the requests. This handles mixed hardware, which is common as estates grow by addition. Least connections sends each request to the server currently handling the fewest. Better than round robin when request durations vary widely, because a server stuck with several long-running requests stops receiving new ones automatically. Weighted least connections ...

LACP and Link Aggregation: Why Four Links Are Not Four Times the Speed

Link aggregation combines several physical links into one logical link, adding bandwidth and redundancy at the same time. LACP is the standard protocol that negotiates it. The single most important thing to understand is what aggregation does not give you: a single conversation does not get faster. Why a flow cannot be split TCP requires packets to arrive in order, or close enough that reordering does not trigger retransmission. If a switch sprayed one session's packets across four links with different queue depths, they would arrive out of order and performance would collapse. So the switch hashes some combination of header fields — source and destination MAC addresses, IP addresses, and often port numbers — and uses the result to pick one link. All packets in a given flow take the same link , arriving in order. The consequence: four bundled 1 Gbps links give 4 Gbps of aggregate capacity across many flows, and any single transfer is still capped at 1 Gbps. A question ...