What is RPO?

RPO (Recovery Point Objective) reflects the data loss in the event of a failure.

If you are looking for a high availability cluster with automatic failover, then the RPO should be 0. The application is thus restarted without data loss. Either you can choose a hardware high availability cluster with shared disk. Or you can choose a software high availability cluster with synchronous real-time replication to have 0 data loss.

If you are implementing backup solutions, then the RPO is greater than 0 and the recovery is not automatic. Administrators decide how often to replicate and how many backups to keep.

What is RTO?

RTO (Recovery Time Objective) is the time during which an application is unavailable in the event of a failure.

For a critical application, RTO should be minimal. For this, a high availability solution is necessary with automatic restart of the application in the event of hardware or software failures. RTO is then approximatively one minute: the detection time plus the automatic restart time of the application.

With a backup solution, RTO is generally greater than several hours. Administrators will first attempt to repair the hardware and restart the application on up-to-date data. Restarting from a backup is the last decision when previous actions don't work, because it leads to data loss.

RTO with the example of a SafeKit mirror cluster

The SafeKit mirror cluster is a software high availability cluster with synchronous real-time data replication and automatic application failover.

RTO of the SafeKit mirror cluster is in the order of 1 mn and can be decreased if you configure the heartbeat timeout.

For a hardware failure, RTO = heartbeat timeout (default 30 s) + time to restart the application.

For a software failure or an administrator restart, RTO = time to stop the application + time to restart it.

With solutions that reboot a full virtual machine in case of failure, the RTO includes the reboot time of the virtual machine.

RTO with the example of a SafeKit farm cluster

The SafeKit farm cluster is a software high availability cluster with network load balancing and automatic failover.

RTO of a SafeKit farm cluster is in the order of a few seconds.

For a hardware failure, RTO = failure detection timeout through monitoring channels (default a few seconds). After the timeout the load balancing filters are reconfigured.

For a software failure or an administrator restart, RTO = time to stop the application + time to restart it.

What are the advantages of a mirror cluster?

Low Complexity
Plug&Play deployment with no specific skills
Suitable for large deployments in many sites (very simple to deploy)
2 physical or virtual nodes
No shared storage requirement
No Domain Controller requirement
Same solution on Windows and Linux
Support Windows Server and Client OS editions
Well documented API and support
Synchronous data replication (no data loss in case of failure)
Replicated directories can be in the system disk
Supports multiple heartbeats and vitual IP addresses
Offers configurable software, hardware and network checkers
For the split brain problem and the quorum, does not require a special disk or a third machine or a dedicated link between both servers
Automatic failover of application with a recovery time in the order of one minute
Automatic failback when a server comes back after a failure (no manual operation)
A very simple console to deploy the solution and to maintain it afterwards for end-customer
Supports hardware and environment failures (20% of causes of unavailability), including the complete failure of a computer room with 2 nodes in two remote sites
Supports software failures (40% of causes of unavailability): software bug, regression on software update (N and N+1 versions can coexist)
Supports human errors (40% of causes of unavailability) : the simplicity of use avoids the administration error of the critical application

What are the advantages of a farm cluster

Low Complexity
Plug&Play deployment with no specific skills
Suitable for large deployments in many sites (very simple to deploy)
2 physical or virtual nodes or more
No network load balancers requirement
No proxy server requirement (above the farm cluster)
No Domain Controller requirement
No restriction in VMware due to multicast or unicast address
Same solution on Windows and Linux
Support Windows Server and Client OS editions
Well documented API and support
Supports multiple monitoring channels on multiple networks for server failure detection
Supports multiple vitual IP addresses
Offers configurable software, hardware and network checkers
Offers the mirror cluster with synchronous real-time replication and failover to implement a farm+mirror 3-tiers architecture
Automatic failover with a recovery time in the order of a few seconds
Automatic failback when a server comes back after a failure (no manual operation)
A very simple console to deploy the solution and to maintain it afterwards for end-customer
Supports hardware and environment failures (20% of causes of unavailability), including the complete failure of a computer room with 2 nodes in two remote sites
Supports software failures (40% of causes of unavailability): software bug, regression on software update (N and N+1 versions can coexist)
Supports human errors (40% of causes of unavailability): the simplicity of use avoids the administration error of the critical application

SafeKit Solutions and Quick Installation Guides

Key differentiators of high availability at the virtual machine level or at the application level

Key differentiators of SafeKit vs Microsoft Hyper-V cluster and VMware HA

Key differentiators of a mirror cluster with replication and failover

Evidian SafeKit mirror cluster with real-time file replication and failover
3 products in 1 More info >	The SafeKit high availability software saves on Windows and Linux the cost of : external shared or replicated storage, load balancing boxes, enterprise editions of OS and databases SafeKit includes all clustering features: synchronous real-time file replication, monitoring of server / network / software failures, automatic application restart, virtual IP address switched in case of failure to reroute clients
Very simple configuration More info >	The cluster configuration is very simple and made by means of application modules. New services and new replicated directories can be added to an existing application module to complete a high availability solution All the configuration of clusters is made using a simple centralized web administration console There is no domain controller or active directory to configure as with Microsoft cluster
Synchronous replication More info >	The real-time replication is synchronous with no data loss on failure This is not the case with asynchronous replication
Fully automated failback More info >	After a failure when a server reboots, the replication failback procedure is fully automatic and the failed server reintegrates the cluster without stopping the application on the only remaining server This is not the case with most replication solutions particularly with replication at the database level. Manual operations are required for resynchronizing a failed server. The application may even be stopped on the only remaining server during the resynchonization of the failed server
Replication of any type of data More info >	The replication is working for databases but also for any files which shall be replicated This not the case for replication at the database level
File replication vs disk replication More info >	The replication is based on file directories that can be located anywhere (even in the system disk) This is not the case with disk replication where special application configuration must be made to put the application data in a special disk
File replication vs shared disk More info >	The servers can be put in two remote sites This is not the case with shared disk solutions
Remote sites and virtual IP address More info >	All SafeKit clustering features are working for 2 servers in remote sites. Replication requires an extended LAN type network (latency = performance of synchronous replication, bandwidth = performance of resynchronization after failure). If both servers are connected to the same IP network through an extended LAN between two remote sites, the virtual IP address of SafeKit is working with rerouting at level 2 If both servers are connected to two different IP networks between two remote sites, the virtual IP address can be configured at the level of a load balancer with the "healh check" of SafeKit.
Quorum and split brain More info >	The solution works with only 2 servers and for the quorum (network isolation between both sites), a simple split brain checker to a router is offered to support a single execution of the critical application This is not the case for most clustering solutions where a 3^rd server is required for the quorum
Active/active cluster More info >	The secondary server is not dedicated to the restart of the primary server. The cluster can be active-active by running 2 different mirror modules This is not the case with a fault-tolerant system where the secondary is dedicated to the execution of the same application synchronized at the instruction level
Uniform high availability solution More info >	SafeKit implements a mirror cluster with replication and failover. But it imlements also a farm cluster with load balancing and failover. Thus a N-tiers architecture can be made highly available and load balanced with the same solution on Windows and Linux (same installation, configuration, administration with the SafeKit console or with the command line interface). This is unique on the market This is not the case with an architecture mixing different technologies for load balancing, replication and failover
RTO / RPO More info >	SafeKit implements quick application restart in case of failure: around 1 mn or less Quick application restart is not ensured with full virtual machines replication. In case of hypervisor failure, a full VM must be rebooted on a new hypervisor with a recovery time depending on the OS reboot as with VMware HA or Hyper-V cluster

Key differentiators of a farm cluster with load balancing and failover

Evidian SafeKit farm cluster with load balancing and failover
No load balancer or dedicated proxy servers or special multicast Ethernet address More info >	The solution does not require load balancers or dedicated proxy servers above the farm for imlementing load balancing. SafeKit is installed directly on the application servers in the farm. The load balancing is based on a standard virtual IP address/Ethernet MAC address and is working with physical servers or virtual machines on Windows and Linux without special network configuration This is not the case with network load balancers This is not the case with dedicated proxies on Linux This is not the case with a specific multicast Ethernet address on Windows
All clustering features More info >	The solution includes all clustering features: virtual IP address, load balancing on client IP address or on sessions, monitoring of server / network / software failures, automatic application restart with a quick revovery time and a replication option with a mirror module This is not the case with other load balancing solutions. They are able to make load balancing but they do not include a full clustering solution with restart scripts and automatic application restart in case of failure. They do not offer a replication option The cluster configuration is very simple and made by means of application modules. There is no domain controller or active directory to configure on Windows. The solution works on Windows and Linux
Remote sites and virtual IP address More info >	If servers are connected to the same IP network through an extended LAN between remote sites, the virtual IP address of SafeKit is working with load balancing at level 2 If servers are connected to different IP networks between remote sites, the virtual IP address can be configured at the level of a load balancer with the help of the SafeKit health check. Thus you can implement load balancing but also all the clustering features of SafeKit, in particular monitoring and automatic recovery of the critical application on application servers
Uniform high availability solution More info >	SafeKit imlements a farm cluster with load balancing and failover. But it implements also a mirror cluster with replication and failover. Thus a N-tiers architecture can be made highly available and load balanced with the same solution on Windows and Linux (same installation, configuration, administration with the SafeKit console or with the command line interface). This is unique on the market This is not the case with an architecture mixing different technologies for load balancing, replication and failover

Key differentiators of the SafeKit high availability technology

Software clustering vs hardware clustering More info >
A simple software cluster with the SafeKit package just installed on two servers	Complex hardware clustering with external storage or network load balancers
Shared nothing vs a shared disk cluster More info >
SafeKit is a shared-nothing cluster: easy to deploy even in remote sites	A shared disk cluster is complex to deploy
Application High Availability vs Full Virtual Machine High Availability More info >
Application HA supports hardware failure and software failure with application checkers. Quick recovery time by restarting only the application (RTO around 1 mn or less). Application HA requires to define restart scripts per application and folders to replicate (SafeKit application modules).	Full virtual machines HA supports hardware failure and some software failures like a frozen VM. VM reboot on failure and recovery time depending on the OS reboot. No restart scripts to define with full virtual machines HA (SafeKit hyperv.safe or kvm.safe modules). Hypervisors are active/active with just multiple virtual machines.
High availability vs fault tolerance More info >
No dedicated server with SafeKit. Each server can be the failover server of the other one.Software failure with restart in another OS environment. Smooth upgrade of application and OS possible server by server (version N and N+1 can coexist)	Secondary server dedicated to the execution of the same application synchronized at the instruction level.Software exception on both servers at the same time. Smooth upgrade not possible
Synchronous replication vs asynchronous replication More info >
SafeKit implements real-time synchronous replication with no data loss in case of failure	With asynchronous replication, there is data loss on failure
Byte-level file replication vs block-level disk replication More info >
SafeKit implements real-time byte-level file replication and is simply configured with application directories to replicate even in the system disk	Block-level disk replication is complex to configure and requires to put application data in a special disk
Heartbeat, failover and quorum to avoid 2 master nodes More info >
To avoid 2 masters, SafeKit proposes a simple split brain checker configured on a router	To avoid 2 masters, other clusters require a complex configuration with a third machine, a special quorum disk, a special interconnect
Virtual IP address primary/secondary, network load balancing, failover More info >
No dedicated proxy servers and no special network configuration are required in a SafeKit cluster for virtual IP addresses	Special network configuration is required in other clusters for virtual IP addresses. Note that SafeKit offers a health check adapted to load balancers

VM HA with the SafeKit Hyper-V or KVM module	Application HA with SafeKit application modules

SafeKit inside 2 hypervisors: replication and failover of full VM	SafeKit inside 2 virtual or physical machines: replication and failover at application level
Replicates more data (App+OS)	Replicates only application data
Reboot of VM on hypervisor 2 if hypervisor 1 crashes Recovery time depending on the OS reboot VM checker and failover (Virtual Machine is unresponsive, has crashed, or stopped working)	Quick recovery time with restart of App on OS2 if crash of server 1 Around 1 mn or less (see RTO/RPO here) Application checker and software failover
Generic solution for any application / OS	Restart scripts to be written in application modules
Works with Windows/Hyper-V and Linux/KVM but not with VMware	Platform agnostic, works with physical or virtual machines, cloud infrastructure and any hypervisor including VMware

SafeKit with the Hyper-V module or the KVM module	Microsoft Hyper-V Cluster & VMware HA

No shared disk - synchronous real-time replication instead with no data loss	Shared disk and specific extenal bay of disk
Remote sites = no SAN for replication	Remote sites = replicated bays of disk across a SAN
No specific IT skill to configure the system (with hyperv.safe and kvm.safe)	Specific IT skills to configure the system
Note that the Hyper-V/SafeKit and KVM/SafeKit solutions are limited to replication and failover of 32 VMs.	Note that the Hyper-V built-in replication does not qualify as a high availability solution. This is because the replication is asynchronous, which can result in data loss during failures, and it lacks automatic failover and failback capabilities.

What is RPO and RTO with examples?

Evidian SafeKit

What is RPO and RTO with examples of high availability and backup solutions?

Overview

What is RPO?

What is RTO?

RTO with the example of a SafeKit mirror cluster

RTO with the example of a SafeKit farm cluster

RPO with the example of a SafeKit mirror cluster

RPO with the example of a SafeKit farm cluster

What are the advantages of a mirror cluster?

What are the advantages of a farm cluster

SafeKit Solutions and Quick Installation Guides

SafeKit High Availability Differentiators

Evidian SafeKit mirror cluster with real-time file replication and failover

Evidian SafeKit farm cluster with load balancing and failover

Software clustering vs hardware clustering More info >

Shared nothing vs a shared disk cluster More info >

Application High Availability vs Full Virtual Machine High Availability More info >

High availability vs fault tolerance More info >

Synchronous replication vs asynchronous replication More info >

Byte-level file replication vs block-level disk replication More info >

Heartbeat, failover and quorum to avoid 2 master nodes More info >

Virtual IP address primary/secondary, network load balancing, failover More info >