Managed IT
Disaster recovery when there are 180 covers booked tonight
Business continuity for hotels and restaurants with covers booked tonight. RTO and RPO explained, the failure scenarios that happen and what restores service.
Disaster recovery for hospitality is judged by whether you can still take a booking, seat a guest and take payment. MicroNet Global plans it around service, not servers: recovery time and recovery point targets per system, tested failover connectivity, rehearsed offline procedures and backups that have been restored, not merely scheduled.
What are RTO and RPO?
Recovery Time Objective is how long a system may be unavailable before the damage is unacceptable. Recovery Point Objective is how much data you can afford to lose, measured in time. In hospitality the two differ sharply by system: an EPOS till has an RTO of minutes and an RPO of near zero, while a back-office reporting server might tolerate a day of downtime and a night's data loss.
Stated in plain operational terms, RTO is "how long until we can serve again" and RPO is "how many bookings, covers or transactions we would have to re-enter by hand". Writing both down per system is the single most useful hour a hospitality operator can spend on continuity planning, because it turns an abstract fear into a shopping list.
The failure scenarios that actually happen
Most hospitality outages are not fires. They are a broadband circuit cut by a contractor two streets away, a supplier's cloud platform going dark at lunchtime, or a ransomware incident that spreads through a flat network overnight.
| Scenario | Typical RTO target | What actually |
|---|---|---|
| Single-site broadband loss | Minutes | Tested 4G/5G failover on the firewall, with payment and PMS traffic prioritised |
| EPOS or PMS cloud vendor outage | Vendor-controlled | Rehearsed offline mode, printed fallback procedures, supplier escalation contacts |
| Ransomware across the group | 24--72 hours | Immutable offsite backups, segmentation that limited spread, clean rebuild capability |
| Power failure at one site | Minutes to hours | UPS on comms and payment kit, safe shutdown, generator if the site warrants it |
| Flood or fire closing a site | Days to weeks | Cloud-hosted core systems, spare kit pool, relocating bookings to sister sites |
| Loss of a back-office server | 4--24 hours | Virtualised workload restored to alternative hardware or cloud; tested restores |
| Key SaaS supplier failure | Vendor-controlled | Contractual RTO commitments, status monitoring, a documented manual workaround |
Protective, detective and remedial controls
A continuity plan is easier to write, and easier to audit, when it is grouped by what each control is for.
**Protective controls** reduce the chance of the outage: segmentation, patching, multi-factor authentication, resilient power, dual circuits, least-privilege access. These are the cheapest controls per pound of risk removed, and the ones most often deferred during a fit-out.
**Detective controls** shorten the time to know: monitoring and alerting on circuits, servers and backup jobs, endpoint detection, and someone watching the alerts at 2am. An outage discovered by a duty manager at service is already an hour old.
**Remedial controls** restore service: failover connectivity, immutable backups, spare hardware, documented rebuild steps and an escalation path with named authority. Our [[24/7 managed IT support]{.underline}](about:blank) sits across all three, because detection without response capability just produces better-documented outages.
The five things operators consistently under-plan
**Failover connectivity that has never been tested.** Many sites have a 4G or 5G backup in the firewall. Far fewer have pulled the primary circuit at 11am on a Tuesday to see whether payments, the PMS and the tills all come back, and whether the failover bandwidth can carry them. An untested failover is a hope, not a control.
**Offline mode nobody has rehearsed.** Most EPOS platforms can trade offline, store transactions and sync later. If the team has never practised it, they will discover the limits during service, and the limits matter: how long it holds, which payment methods still work, and what happens when the queue syncs.
**Printed fallback procedures.** When the network is down, the procedure stored on SharePoint is unavailable. One laminated sheet per site --- the offline steps, the escalation numbers, the manual card process --- has saved more services than any piece of software.
**Out-of-hours escalation and the authority to declare.** At 9pm on a Saturday, someone has to decide to close bookings, switch to manual or call the supplier's emergency line. Name that person per site and per group, give them a deputy, and make sure the IT partner's escalation path is tested rather than assumed.
**Backups that are tested, not scheduled.** A backup job with a green tick is not a restore. Follow 3-2-1 --- three copies, two media types, one offsite --- and add at least one immutable copy that ransomware cannot encrypt or delete. Then restore something, on a schedule, and record the time it took. That measured figure is your real RTO; everything else is an estimate.
One site is an incident. Twenty sites is a different problem.
Group-level continuity is not site continuity repeated. If the same incident hits twenty venues at once, you do not have twenty engineers, twenty spare firewalls or twenty people free to answer the phone, and the finite resource becomes coordination rather than technology.
That changes the plan in three ways. You need a triage order agreed in advance --- which sites recover first, usually by covers, revenue or contracted events. You need a single incident channel, so twenty general managers are not calling one engineer. And you need standardised builds, because a group where every site was configured differently cannot be recovered in parallel.
Standardisation is also what makes remote recovery possible at all. Consistent hardware, documented configurations and central management are why around 90% of the network faults MicroNet Global handles are resolved remotely, with no site visit. That ratio collapses in an estate assembled site by site, a common finding when groups grow by acquisition --- the trade-offs are covered in our comparison of an [[in-house IT team versus a hospitality IT partner]{.underline}](about:blank).
How MicroNet Global approaches it
MicroNet Global has specialised in hospitality since 2004 and supports 745+ sites across 15+ countries from offices in the UK, US and UAE. Continuity work starts with a per-system RTO and RPO register agreed with operations, not with a product.
From there it is [[IT infrastructure design and disaster recovery]{.underline}](about:blank) --- resilient connectivity, segmentation, backup architecture with immutable copies --- and then the unglamorous part: scheduled restore tests, failover tests booked into quiet trading periods, and fallback documentation kept current. Incidents run through a 24/7 global service desk with a 4-hour remote resolution target and an 8-hour onsite target, and the same discipline applies when systems break at the worst possible moment, as in [[hospitality cyber security attack paths]{.underline}](about:blank). If you want a continuity review of your estate, [[talk to the team]{.underline}](about:blank).
Frequently asked questions
Minutes, not hours --- tills are the last system you can trade without. In practice that means the recovery mechanism is offline mode and failover connectivity rather than a restore: the terminals keep trading locally while the circuit or the platform comes back. Set the RPO near zero and confirm with your EPOS vendor exactly how long offline mode holds before it stops accepting transactions.
Written by the MicroNet Global team. If you are working through any of this for your own estate, the specialists here are happy to talk it through.
