Cloud & Architecture 8 min read Quarterly Strategic Edition

Running SAP Workloads on Cloud: Architecting for 99.99% Availability on Azure and AWS

A comparative architectural guide evaluating high-availability clusters, storage throughput with Azure NetApp Files vs AWS EBS, and disaster recovery SLA guarantees.

DT

Cloud Infrastructure Practice

Pixelverge LLC Strategic Research Group

Target: Cloud Infrastructure Directors and Solutions Architects
Executive Overview

SAP workloads are different from most cloud workloads. They demand sustained memory bandwidth, low-latency storage, strict high-availability patterns, and disaster recovery that the business treats as a promise rather than a nice-to-have. The good news is that both Azure and AWS have mature, certified paths for running SAP. The challenge is choosing—and architecting—the right one. This guide compares the two hyperscalers on the dimensions that actually matter for SAP, and lays out the architecture patterns that get enterprises to 99.99% availability without guessing.

At a Glance
  • Multi-zone HANA System Replication (HSR) with automated Pacemaker fencing
  • Storage IOPS considerations for large-scale HANA in-memory memory loading
  • Cost optimization using 3-year Reserved Instances and automated non-production scaling
  • Meeting RPO < 15min and RTO < 2hr across geographic disaster recovery pairs
1

Why SAP Cloud Architecture Is Different

A web application can tolerate brief outages. An ERP cannot. When SAP is down, order entry stops, inventory freezes, and finance closes late. That is why SAP cloud architecture is built around availability zones, system replication, and deliberate failure design rather than simple scale-out.

Both Azure and AWS are certified for SAP—including S/4HANA and SAP HANA—with published instance families, validated reference architectures, and supported high-availability patterns. The choice between them is rarely about capability and usually about fit: how the rest of your estate runs, your identity and security model, and where your teams already have depth.

Whatever the platform, the architecture fundamentals are the same: certified instances, replicated databases, automated failover, and tested disaster recovery.

  • ERP availability is a business requirement, not a technical preference
  • Both hyperscalers offer SAP-certified instances and validated patterns
  • Choice comes down to estate fit, identity, and team depth

The platform is the easy decision. The architecture is where availability is won or lost.

2

High Availability: Multi-Zone Clustering

High availability for SAP HANA means HANA System Replication (HSR) between instances in different availability zones, with automated failover coordinated by a cluster manager—Pacemaker on both platforms—including fencing to prevent split-brain scenarios.

On Azure, this maps to availability zones with proximity placement groups, ensuring the primary and secondary HANA instances sit close enough for synchronous replication while remaining isolated from zone-level failures. On AWS, Multi-AZ deployments with the same HSR pattern give equivalent protection.

The goal is automated, fast failover that the business does not have to think about. Manual failover is not high availability; it is a longer incident.

  • HANA System Replication across availability zones
  • Pacemaker cluster with fencing on both Azure and AWS
  • Automated failover the business does not have to think about
3

Storage: Throughput Is the Hidden Constraint

SAP HANA is an in-memory database, but it is not storage-light. The database must load into memory on restart, log writes continuously, and support snapshots and replication. Under-provisioned storage shows up as slow restarts, checkpoint stalls, and replication lag.

On Azure, Azure NetApp Files provides the high-IOPS, low-latency file storage SAP HANA expects. On AWS, the equivalent is high-performance EBS volumes—io2 Block Express for the database, backed by the right throughput and burst profiles for the workload.

Storage sizing is not a one-time calculation. It depends on the HANA memory footprint, transaction volume, and recovery-time targets. Getting it wrong is one of the most common causes of “the architecture was right but the performance was not.”

  • Match storage IOPS and throughput to the HANA memory footprint
  • Azure NetApp Files vs AWS EBS io2—choose for sustained throughput, not peak burst
  • Design for fast restart and replication lag, not just steady-state load

In-memory does not mean storage doesn’t matter. It means storage matters differently.

4

Disaster Recovery: RPO and RTO as Promises

Disaster recovery for SAP is defined by two numbers: recovery point objective (RPO)—how much data you can afford to lose—and recovery time objective (RTO)—how quickly you must be back. For most enterprises, RPO under 15 minutes and RTO under 2 hours are realistic and testable targets.

Both platforms support async or sync HANA System Replication to a secondary region, automated DNS failover, and orchestrated failover runbooks. The difference between a DR plan and a DR promise is rehearsal. An un-rehearsed failover is a hypothesis.

The architecture should include infrastructure-as-code so the recovery environment can be stood up predictably, and automated DR validation so the plan is tested on a schedule rather than discovered to be broken during an actual event.

  • Define RPO and RTO as business commitments, then architect to meet them
  • Replicate to a secondary region with automated failover runbooks
  • Rehearse DR on a schedule—an untested failover is a hypothesis

An untested disaster recovery plan is a plan for the disaster to be worse.

5

Cost: Architecting for Predictability

SAP workloads are expensive enough that cost discipline is an architectural concern. Reserved instances or savings plans on a 3-year horizon materially reduce the bill for always-on production, while non-production environments can scale down automatically outside business hours.

The temptation is to over-provision “just in case.” The discipline is right-sizing against measured utilization, with headroom governed by the availability targets rather than by fear.

A FinOps layer—tagging, budgets, and automated non-production scaling—keeps the architecture honest as the estate evolves.

  • Use 3-year Reserved Instances or Savings Plans for always-on production
  • Automate non-production scale-down outside business hours
  • Right-size against measured utilization, not worst-case fear
Closing Perspective

The hyperscaler you choose for SAP matters less than the architecture you build on it. High availability, storage throughput, tested disaster recovery, and cost predictability are the four pillars that determine whether the workload meets the business’s expectations. Both Azure and AWS can host world-class SAP landscapes—the difference is in the design, the rehearsals, and the discipline that keeps the architecture true over time. If you are planning a cloud migration or re-platforming an existing landscape, start from the availability and recovery targets the business actually needs, and let those drive every architectural decision.

Frequently Asked Questions

Both are certified and capable for SAP. The choice depends on your broader cloud estate, identity and security model, integration needs with Microsoft 365 or the AWS ecosystem, and where your teams already have operational depth.

Continue Reading

Executive Strategy

Clean Core Strategy: Extending SAP Without Building Tomorrow’s Technical Debt

How leading enterprises keep their S/4HANA core pristine using SAP BTP, side-by-side extensions, and developer extensibility to ensure effortless future upgrades.

Read Publication
Checklist & Guides

Enterprise S/4HANA Readiness Checklist: 25 Pre-Flight Checks for IT Leaders

A structured evaluation framework covering SAP landscape sizing, data archiving, simplification items, interface mapping, and testing governance.

Read Publication
EXECUTIVE DIALOGUE

Discuss This Framework With Our Practice Leads

Schedule a confidential 30-minute peer review of your transformation roadmap.

Email Us Consult