
Principal SRE
At a Glance
- Category
- 💻 Technology
- Level
- Senior
- Type
- Full-time
Spot the Problem
- Find what's costing you interviews at Alpheya
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
About Alpheya
Alpheya is a wealth-management technology company headquartered in Abu Dhabi. Banks are our customers. We build each one a tailored investing experience, from mobile apps to advisor and back-office portals, on a single shared platform covering the full order-to-custody lifecycle. The platform runs as SaaS on Microsoft Azure, with on-premises delivery for banks that require it.
The role
Over the next six months we are taking multiple banks live as SaaS customers. Our engineering teams build the platform and our SREs run the infrastructure. What we don't yet have is a single leader accountable for the service our customers buy: the SLAs, the incident review a bank CIO sits in on, the audits, the disaster-recovery program, the cost of running each tenant. That is this role.
You will report to the CTO and own SaaS operations end to end, from the Kubernetes clusters to the quarterly service review with a bank's executives. The mandate: make onboarding the fifth bank a checklist instead of a project.
What you'll own
- Service management for every SaaS customer. SLA definition and reporting, incident communications and post incident reviews, security questionnaires, audit cycles, and the day-to-day support model with each bank's service desk.
- Release and deployment operations. The release calendar across the tenant estate, environment promotion and rollback discipline, production change management, and coordinating rollouts with each bank's change and freeze windows. Engineering builds the pipelines; you decide when and how production changes.
- The reliability program. Tested disaster recovery and business continuity, backup verification, patching and vulnerability-management cadence, capacity planning, and a cost-per-tenant model the CFO can use.
- The tenant onboarding runbook, so new banks go live on a repeatable path.
- Our European launch. Two new Azure regions with data residency, DR, and support coverage in place, operated rather than merely deployed.
- Compliance operations for outsourcing arrangements, working with bank risk teams under frameworks such as ISO 27001, SOC 2, and European and Gulf outsourcing regulation.Your first six months
- Four bank go-lives across two geographies, each with agreed SLAs, escalation paths, and incident procedures in place from day one.
- A disaster-recovery exercise run and documented for at least one production environment.
- An on-call rotation covering both regions without heroics.
- A tenant cost model and a capacity plan for the whole estate.
- 12+ years in production operations, several of them leading the function for multi-tenant SaaS.
- Ownership of customer-facing service management. You have led a severity-one bridge, presented the post-incident review to a customer's executives, and been through their audits.
- You have built or scaled an SRE or platform-operations team and designed on-call across regions.
- Enough Kubernetes and cloud depth to challenge your engineers on failure modes, DR design, and capacity claims.
- Working knowledge of Azure itself. Regions and availability zones, networking and private connectivity, identity, and quota planning. You will negotiate these with Microsoft and with bank security teams.
- Experience with certification and audit regimes: ISO 27001, SOC 2, or bank outsourcing regulation in the EU or the Gulf.Nice to have
- Wealth management, brokerage, or capital-markets domain exposure.
- Experience operating streaming or workflow infrastructure (Temporal, Kafka, or similar).
- Experience with on-premises software delivery to enterprise customers.
Requirements
- •12+ years in production operations
- •Experience leading production operations for multi-tenant SaaS
- •Experience leading severity-one incident bridges and presenting post-incident reviews to executives
- •Experience building or scaling SRE or platform-operations teams
- •Experience designing on-call rotations across regions
- •Deep knowledge of Kubernetes and cloud infrastructure
- •Experience with disaster recovery design and capacity planning
- •Knowledge of ISO 27001, SOC 2, and outsourcing regulations
Responsibilities
- •Own SaaS operations end-to-end, including SLAs, incident reviews, audits, and disaster recovery
- •Manage service management for every SaaS customer
- •Define and report SLAs, manage incident communications, and handle security questionnaires
- •Oversee release and deployment operations, including production change management
- •Lead the reliability program (backup verification, patching, vulnerability management)
- •Develop a tenant cost model and capacity plan for the entire estate
- •Create a repeatable tenant onboarding runbook
- •Manage European launch operations across new Azure regions
Related Jobs2 similar jobs
Browse Similar
- Find what's costing you interviews at Alpheya
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
Alpheya offers digital transformation and cloud solutions, helping businesses modernize their operations and leverage technology for growth. They serve various industries.

