Skip to main content
Versa Networks

Director High Availability Architecture

Versa Director supports a clustered deployment model that provides high availability (HA) and automatic failover for mission-critical management of Versa Operating SystemTM (VOSTM) appliances. A Director cluster eliminates single points of failure, ensuring continuous network management even when individual nodes experience outages.

The clustered deployment model provides the following benefits:

  • Automatic failover—Provides automatic failover during unplanned node failures with no manual intervention required.
  • 3-node cluster design—Employs Apache ZooKeeper® for quorum-based decision-making, ensuring failover proceeds only when a majority of nodes agree. This eliminates the split-brain problem inherent in 2-node designs, where each node believes the other has failed and begins operating independently, which leads to data inconsistency or corruption.
  • Synchronous database replication—Writes configuration and state changes to database replicas before confirming success, so acknowledged changes are not lost during failover.
  • Encrypted inter-node communication (IPsec full mesh)—Establishes direct, encrypted IPsec tunnels between every node pair, forming a fully-connected mesh that encrypts all inter-node traffic in transit, including replication, quorum, and data traffic, to protect against eavesdropping and tampering.
  • Shared data durability—Uses GlusterFS 3-way replicated storage for configuration, certificates, and packages, providing HA and fault tolerance by allowing the cluster to survive node failures while keeping data accessible, and improving read performance by distributing requests across all three replicas.
  • Parallel device commits—Processes commits to different devices in parallel, unlike the serial commit design used by some systems. This improves overall throughput and provides the scalability and isolation required for a true multi-tenant architecture.
  • Horizontal scalability—Supports horizontal scaling, enabling a single deployment to manage a large number of devices while providing a unified UI for provisioning, monitoring, and management.

Cluster Topology

The recommended production deployment uses a 3-node cluster. Each node has a distinct role: primary, secondary, or arbiter.

The primary node is the leader of the cluster, with the authority to run core services and manage devices. The leadership cannot change unless a majority of cluster nodes agree, which is referred to as a quorum. In a 3-node cluster, maintaining a quorum means at least two nodes must remain connected for failover or promotion to proceed. If the quorum is lost, the cluster freezes state transitions to prevent split-brain (two nodes claiming primary at once).

Roles

Assigning a distinct role to each node in the cluster ensures that application load, secondary capacity, and quorum voting are separated.

Only one node acts as the primary at any given time. The secondary node stands by with a synchronized database replica, ready to assume the primary role upon failover or switchover. The arbiter node does not run core services. Its primary purpose is to hold the third vote in quorum decisions and prevent split-brain scenarios. If either the primary or secondary node becomes unreachable, the surviving node pairs with the arbiter to form a 2-of-3 majority, allowing the cluster to confidently elect or maintain leadership.

The table below summarizes the roles for each type of node. 

Role  Purpose
Primary 

Runs core services and common services, and manages Versa VOS appliances.

Secondary

Keeps common services and a synchronous database replica ready, but does not run core services until failover or switchover. Takes over as primary if the current primary node becomes unreachable.

Arbiter

Quorum voter and lightweight participant. Runs common services needed for consensus (including PostgreSQL with no failover), but does not host core services. Provides the quorum third vote so the cluster can safely decide leadership without split-brain.

Services and Orchestration

Director runs its services as Docker Swarm stacks across all three nodes. Swarm operates over an overlay network that spans either the northbound or the southbound network.

From the Docker Swarm manager, you assign each node's role—primary, secondary, or arbiter—as a Docker Swarm node label. These labels control where Swarm schedules core services. Placement constraints in the service definitions restrict all core services to the node acting as primary, including the Web UI, Backend API, UniConfig, Redis, OpenBao, NetBox, and related components.

During failover, Director swaps the labels, relabeling the previous primary as secondary and promoting the selected secondary to primary. Swarm then automatically migrates core services to the new primary. The arbiter keeps its label throughout failover. That node continues serving as a quorum voter and never hosts core services.

Traefik, which runs as a global Swarm service, load balances northbound traffic for the cluster. Unlike the core services, which run only on the primary node, Traefik runs on every node, terminating northbound HTTPS/REST traffic and routing it to the active services on the primary. You can verify cluster-wide service health by running the vsh status command from the shell to view status across the entire cluster rather than on a per-node basis.

The table below summarizes the service functions and behaviors. 

Service

Function HA Behavior
Traefik Load balancer, HTTPS termination, routing

Runs on every node (global). Routes northbound traffic only to backends on the current primary node. 

Web UI

Director management GUI

Runs on the primary node. 
Backend API Core REST API for all management operations Runs on the primary node. 
UniConfig Southbound device controller (Netconf) Runs on the primary node. 
PostgreSQL (Patroni) Configuration and state database Synchronous replication and automatic leader election.
Apache ZooKeeper Distributed coordination and consensus Runs on every node as an ensemble. Requires a majority (2-of-3) quorum for consensus. 
Apache Kafka ® Event streaming and messaging Broker on every node. Topic replication. 
Redis
 

Redis is a fast, temporary memory store that lives on the primary node. Redis stores information such as active user sessions and cached appliance state.

The Redis cache is rebuilt from the PostgreSQL database on failover.
OpenBao Secure storage and management of credentials, certificates, and encryption keys Runs on the primary node. Data is synced to the secondary node. 
NetBox IP address management (IPAM) Runs on the primary node. Backed by a replicated database.
mgmt-service Health monitoring and failover orchestration Runs on every node. Leader-elected. 
node-exporter System metrics collection (CPU, memory, disk) Runs on every node.
versa-host-update Host-level configuration management Runs on every node.

 

Architecture Diagram

ha_cluster_architecture_a4.png

Automatic Failover

During automatic failover, when the primary node becomes unreachable, the cluster promotes the secondary node to become the new primary. All core services start on the new primary, which restores service availability. The system rebalances Netconf notifications to the new primary after failover.

Note: Automatic failover requires that at least two nodes remain connected, forming a quorum majority. This ensures the cluster can safely agree on a new primary without risking split-brain scenarios. 

Automatic Failover Diagram.png

Key behaviors:

  • Fully automatic failover—No operator intervention required.
  • Grace period—The management (mgmt) service polls cluster health every 60 seconds. The secondary node is promoted to primary only after 3 consecutive unhealthy primary checks (approximately 3 minutes), so brief network glitches do not trigger failover.
  • Database promotion—The database is promoted on the secondary node (synchronous replication).
  • Node labels—Node labels are updated so the outgoing primary node is marked secondary, and the promoted node is marked primary. Docker Swarm then reschedules core services onto the new primary.
  • Node rejoining—Once the failed node recovers, it rejoins the cluster as a secondary.
  • Quorum loss—If a quorum is lost (for example, when two nodes go down simultaneously), the cluster blocks failover to prevent corruption.

HA Features

UniConfig is the southbound device controller and the southbound counterpart to Traefik's northbound role. UniConfig maintains southbound Netconf sessions to Versa VOS appliances from the active primary node. This preserves Netconf device management across failover.

After failover:

  • UniConfig and related services restart on the new primary node.
  • Devices reconnect to the management path once the new primary node is ready.
  • The configuration state remains in the replicated database.
  • Netconf notifications are rebalanced to the new primary node after failover or switchover.

Director HA also includes the following features to keep control-plane decisions safe and device management continuous:

  • Patroni PostgreSQL HA—Automatic database leader election independent of which node hosts core services.
  • GlusterFS shared storage—Configuration, certificates, and packages kept consistent across nodes; replication over secure TLS.
  • IPsec mesh—Inter-node traffic encrypted end-to-end.
  • Health monitoring—Failure detection and recovery driven by mgmt-service, node-exporter, and service health checks.
  • Defense in-depth—Multiple health checks, brief delay intervals to disregard transient network disruptions, and a post-failover cooldown to prevent repeated primary role changes.

Data Protection Features

This section describes data protection features of Director HA configurations.

  • Database replication is implemented using PostgreSQL, a relational database management system, and Patroni, its high-availability manager. Three databases run on this replicated infrastructure: vnms (Director configuration), NetBox (IPAM), and UniConfig (device state). 
    • Replication type—Synchronous streaming replication
    • Leader election—Automatic by ZooKeeper consensus
    • WAL archiving—Continuous with LZ4 compression for point-in-time recovery
    • Replica rejoin—Fast incremental sync (pg_rewind) avoids full re-clone
  • Replicated Shared Storage is implemented using GlusterFS, a 3-way replicated distributed file system shared across all nodes, providing fault tolerance and improved read performance by distributing requests across replicas. The file system stores the following:
    • System configuration and cluster state
    • TLS certificates and private keys
    • Software packages and firmware images
    Storage protection features include TLS-encrypted replication, automatic self-healing of file replicas, split-brain detection with majority-based resolution, and quota monitoring with alerts.

Inter-node Security

Director clusters secure all inter-node communication using the following mechanisms:

  • IPsec full mesh—IKEv2 transport-mode tunnel between every pair of nodes
  • Storage encryption—GlusterFS replication over TLS with per-node certificates
  • Firewall hardening—Automated iptables rules restrict cluster-internal ports from external access

Failure Scenarios

The table below summarizes common failure scenarios. In each case, a minimum of two cluster nodes must be reachable and in agreement to designate a primary node. A node that fails or loses contact with the cluster does not remain primary, preventing split-brain conditions.

Scenario What Happens Service Impact How the Cluster Recovers
Primary node failure Primary node becomes unreachable. secondary and arbiter nodes confirm loss of quorum membership for that node. Core services are briefly unavailable while promotion is performed.  Automatic failover to secondary. Former primary rejoins later as secondary.
Secondary node failure Secondary node is lost; primary and arbiter remain connected. No impact to running core services. Cluster continues on primary; secondary rebuilds replicas on recovery.
Arbiter node failure One quorum voter is lost; primary + secondary still form a majority. No impact to running services Quorum maintained; arbiter rejoins when healthy.

Primary isolated

(network

partition)

The former primary is cut off from the other two nodes. It loses quorum and is demoted to secondary on the majority side.

Services cannot be accessed through the isolated primary's Traefik. Majority side continues or promotes secondary to primary.

When connectivity returns, the isolated node re-syncs and rejoins as secondary.

Isolated Primary

If the node that was primary is the one isolated (minority of 1 against a majority of 2):

  • Cluster state—The isolated node loses quorum. ZooKeeper and mgmt-service freeze leadership changes on that side. Patroni cannot maintain an authoritative database leader. Split-brain resolution on the majority side demotes the isolated node to secondary.
  • Traefik—Because the primary is isolated by the network partition, management services cannot be accessed through that node's Traefik. Use the cluster virtual IP address (VIP) or any node on the majority side.
  • On reconnect—The former primary re-syncs PostgreSQL and GlusterFS from the majority side and rejoins as secondary. It does not automatically reclaim primary.

Normal operation diagram.png

Primary Rejoin After Failover

When the original primary recovers it does not automatically reclaim leadership. It reconnects, catches up database and storage state, and rejoins as secondary while the current primary continues serving Versa VOS appliances without interruption.

After Failover Diagram.png

Split-Brain Prevention

All failover decisions require a quorum majority. If the consensus layer (ZooKeeper) does not have a majority, state transitions are frozen.  The system prioritizes correctness over availability to prevent conflicting leader elections.

Note: If a node is partitioned away from the majority, it stops serving management traffic. The majority side continues and may promote the secondary to primary if the isolated node was primary. This prevents split-brain scenarios.

Firewall Requirements 

For the complete list of ports and firewall rules required for Director cluster operation, see Firewall Requirements.

Hardware Requirements

For hardware and software requirements for Versa Director and other headend components, see Hardware and Software Requirements for Headend.

Best Practices

The following table describes best practices for a Director cluster deployment.

Recommendation Details
Distribute nodes across three data centers Deploy the three nodes across three different data centers for site-level redundancy. If a third data center is not available, place the third node on a different host in the secondary data center or in the cloud.
Inter-node latency should be less than or equal to 100 milliseconds Network round-trip latency between cluster nodes must not exceed 100 milliseconds. Higher latency impacts database replication, consensus timeouts, and failover detection.

Cluster interface set to 1500 MTU

The cluster IP interface (the interface used for inter-node cluster communication) on every Director node must be configured with a maximum transmission unit (MTU) of 1500, and the MTU must be consistent across all cluster nodes.
Consistent time synchronization All nodes must be NTP-synchronized. Clock skew can affect certificate validation, log correlation, and consensus mechanisms.
Adequate disk space Ensure sufficient disk on all data nodes for GlusterFS replication, WAL archives, and database storage. Monitor usage with built-in alerts.

Supported Software Information  

Releases 23.1.2 and later support all content described in this article.

  • Was this article helpful?