Proxmox Lab
Flat isometric dark server tower on a glowing pink platform, pink indicator bars down its faces, two pink slab modules wired to opposite corners.
cluster-design

Proxmox VE Cluster Planning: Quorum, Storage, Network

Plan a Proxmox VE cluster with two-node QDevice voting, redundant Corosync links, shared storage, Ceph capacity, and recovery when quorum is lost.

By Proxmox Lab Editorial · ·Updated · 9 min read

A Proxmox VE cluster groups physical hosts under shared configuration and management. Planning it means answering three separate questions: which nodes may make decisions after a failure, where the surviving nodes can read guest disks, and whether they have enough resources to restart the affected guests. Quorum answers only the first question.

Start with the failure you intend to survive: one host, one switch, one storage device, or maintenance on a node. Write that down before choosing a node count. A design can survive a host shutdown and still stop serving applications when its only storage switch fails.

What is a Proxmox cluster?

The cluster manager uses Corosync for membership and distributes configuration through the Proxmox Cluster File System, pmxcfs. You can manage the cluster through a node’s web interface. This does not combine every CPU into one machine, and sharing guest configuration does not copy the guest’s disks.

Distinguish planned migration from high availability. Migration moves a guest between available hosts. HA can restart a managed guest after its previous host has failed and been fenced, provided another eligible node has the necessary storage and resources. Decide which guests require automatic recovery and which can wait for a manual restore.

Use a planning sheet with one row per service: configured memory, disk location, permitted destination nodes, recovery-time target, and acceptable data loss. Those last two entries determine whether scheduled replication is sufficient or whether the service needs shared storage and stronger application-level recovery.

Quorum arithmetic before node count

With the ordinary one-vote-per-node configuration and no QDevice, quorum requires more than half the expected votes. These examples are arithmetic for that configuration, rather than guarantees about application availability.

Cluster nodesVotes requiredNode losses that preserve a majority
220
321
431
532

A fourth node adds compute capacity, but it does not increase the number of failures tolerated by the voting majority compared with three nodes. Conversely, a tiny third node contributes a vote without necessarily providing useful recovery capacity.

The cluster manager documentation describes three nodes as the usual starting point for reliable HA quorum and supports an external vote for smaller two-node arrangements. Count expected votes separately from the machines currently reachable. Powering off a member does not turn the remaining machines into a newly sized cluster.

Two-node cluster plus QDevice

A two-node Proxmox cluster can use an independent arbitrator to supply a third vote. Each Proxmox node runs corosync-qdevice; the external machine runs corosync-qnetd. The external service participates in the voting decision, not in guest execution or disk replication.

In the standard two-node arrangement, both nodes together supply two of three votes even if the arbitrator is unavailable. One node plus the arbitrator’s granted vote can also reach two. A lone node that cannot obtain that external vote cannot form a majority. If the Proxmox nodes lose contact with each other, the arbitrator must grant its vote to only one partition.

The QDevice manual explains the arbitration algorithms and connection behavior. For planning, draw the dependency path to the external server. Hosting it as a guest on one of the two protected nodes creates a dependency on the failure it is supposed to help resolve. Include its power and network path in the diagram.

Two-node voting also leaves a demanding capacity requirement: the survivor must accommodate all guests selected for recovery. A QDevice adds no RAM, CPU, disk copies, or Ceph monitors. It cannot make a two-host Ceph layout equivalent to three independent storage hosts.

Corosync needs predictable latency more than high throughput. Keep its primary path away from bulk migration, backup, and Ceph recovery traffic. A quiet primary link and a physically separate fallback are more useful than two logical networks that disappear with the same cable.

The names ring0_addr and ring1_addr remain in Corosync configuration, while current Proxmox interfaces describe them as Link 0 and Link 1. They identify each node’s address on the corresponding transport link; they do not require a physical ring of cables.

The Corosync configuration manual describes Kronosnet link modes and knet_link_priority. In the passive mode used for Proxmox’s documented fallback arrangement, the highest-priority connected link carries traffic. With equal priorities, the lower link number takes precedence. A second configured link therefore need not carry an equal share of traffic to be useful.

For each link, record the NIC, switch, subnet, MTU, and route between every node. Two VLANs on one NIC still share that NIC’s failure. Two switch ports can still share one switch power supply. Plan a supervised maintenance check for each path separately, with console access and service recovery arrangements ready.

Shared storage options for live migration

The Proxmox storage chapter distinguishes shared storage from node-local storage. Shared storage lets source and destination access the same guest disks, avoiding a full disk transfer during migration.

Storage optionHow hosts access guest disksPlanning question
NFSGuest disk files on an exported filesystemCan every destination reach the export if a storage path fails?
Shared LVM over iSCSILogical volumes on a shared block targetAre target access, LUN presentation, and multipath configured consistently?
Ceph RBDDistributed block images through the Ceph networkDo replicas, monitors, and usable capacity survive the chosen failure?

Use a backend’s supported configuration. Ordinary local LVM-thin does not become a shared backend just because the same storage ID appears on every node. Likewise, marking a directory as shared tells Proxmox about an existing property; it does not replicate files or provide a storage service.

A shared NAS or array also has its own availability requirements. Three compute hosts attached to one unavailable export still have unavailable disks. Include that storage service, its switches, and the maintenance procedure in the recovery plan.

Local disks, live migration, and ZFS replication

Shared storage is helpful, but it is not an absolute requirement for live VM migration. The QEMU/KVM chapter documents transferring local disks to the destination. Allow enough destination space and migration bandwidth, and keep that transfer off the latency-sensitive Corosync path.

VM compatibility matters too. Compatible CPU features, destination software, and available device resources are prerequisites. Physical PCI or USB passthrough can block live migration. Check each guest’s configuration before promising an uninterrupted maintenance move.

Scheduled ZFS replication addresses a different requirement: keeping a disk copy on another node. According to the storage replication chapter, subsequent replication transfers changes, and a migration to a replicated destination can use the existing copy. This reduces the amount left to transfer.

Replication is asynchronous. If a node fails between successful synchronizations, recovery can lose changes made since the last one. Record the acceptable interval per service, monitor failed jobs, and ensure the intended recovery node actually has a usable replica. A scheduled job that has not completed recently provides a different recovery point from the one its schedule implies.

Ceph capacity and recovery planning

Ceph adds a storage cluster with its own health and quorum concerns. Proxmox Corosync votes do not replace Ceph monitor quorum. Plan those systems separately even when their services share physical machines.

The hyper-converged Ceph chapter recommends at least three servers, dedicated storage bandwidth of at least 10 Gbit/s, and memory headroom for recovery. Its OSD guidance recommends approximately 8 GiB per daemon for good performance; four OSDs therefore imply a planning allowance of roughly 32 GiB before guest memory and other host services.

For a replicated pool, size controls the replica count and min_size the minimum replicas needed for I/O. Proxmox’s documented defaults are three and two. Verify the actual pool settings and CRUSH failure domain rather than inferring them from the number of disks.

With three equal storage hosts and replicas placed across hosts, losing one leaves fewer available copies. Spare disk space on the other two hosts cannot recreate three distinct host placements. A replacement host or another eligible failure domain is needed to regain that layout. Budget space for recovery without treating every byte of raw capacity as usable guest storage.

What happens when quorum is lost?

The pmxcfs documentation specifies that configuration becomes read-only without quorum. The UI may return cluster not ready - no quorum? (500); writes under /etc/pve can fail despite a healthy system disk.

Quorum loss alone does not mean every existing non-HA guest immediately stops. A guest may continue executing while management changes are blocked. Disk availability is another condition: an inaccessible NFS export or Ceph pool can interrupt I/O even on a quorate host.

HA changes the failure behavior. The HA manager chapter describes watchdog fencing: an active HA node unable to maintain its required lock can reset itself. Recovery elsewhere follows fencing so the cluster avoids two instances writing the same guest disks.

Treat these as different observations in an incident record: voting status, guest execution, storage health, and fencing state. Record read-only diagnostics such as pvecm status, corosync-cfgtool -s, and ha-manager status before changing configuration. The Proxmox no-quorum troubleshooting guide covers the diagnostic sequence.

Lowering expected votes is an emergency recovery decision after establishing that competing nodes cannot act on the same storage. It is not a substitute for redundant links or a missing third vote.

Size the surviving nodes

Calculate a per-node host reserve, storage-service allowance, and guest allocation. The published hardware requirements start with 2 GB for the OS and Proxmox services, then additional guest memory and storage overhead. The general per-TB storage guideline does not replace a Ceph OSD budget or workload-specific planning.

For a worked planning example, suppose three equal nodes must carry 96 GiB of configured guest RAM after any one node fails. Across two survivors, that is an average of 48 GiB per host before host and storage reserves. A large indivisible guest or a placement restriction can make one node’s requirement higher than that average.

Write down the largest guest, eligible recovery nodes, and services that may stay stopped during an outage. Check CPU capacity and storage throughput alongside memory. Quorum cannot create capacity that was never reserved.

The Proxmox VE HA cluster and Ceph sizer helps explore node counts and its stated RAM assumptions. Independently recalculate the workload on the surviving nodes; a quorum indicator is not proof that every guest will fit.

Maintenance, backups, and migration acceptance

Before relying on the cluster, turn the design into an acceptance checklist. Confirm that each planned migration destination has the right guest network, disk access, and capacity. Record baseline membership and link status while the system is healthy. Schedule recovery exercises according to the service’s tolerated interruption.

Keep backups outside the storage failure you are protecting against. Proxmox Backup Server maintenance guidance describes verification jobs for detecting backup corruption. Pair verification with a restore exercise: recover a guest under a separate ID, isolate its network, and check that its application data is usable. Replication and HA do not recover an earlier version of accidentally deleted data.

For a VMware estate, finish this target-cluster plan before scheduling guest imports. The VMware vSphere to Proxmox VE migration guide covers licensing preparation, import choices, and cutover checks. A successful first boot is one milestone; backup coverage and a workable failure plan are the acceptance criteria that follow.

Sources

  1. Proxmox VE Administration Guide: Cluster Manager
  2. Proxmox VE Administration Guide: Deploy Hyper-Converged Ceph Cluster
  3. Proxmox VE Administration Guide: High Availability
  4. Proxmox VE system requirements
  5. Proxmox VE Administration Guide: Storage
  6. Proxmox VE Administration Guide: QEMU/KVM Virtual Machines
  7. Proxmox VE Administration Guide: Storage Replication
  8. Proxmox VE Administration Guide: Proxmox Cluster File System
  9. Corosync configuration manual
  10. Corosync QDevice manual
  11. Proxmox Backup Server: Maintenance Tasks
#proxmox #clustering #corosync#ceph#high-availability

Related