
TAWANTECH
Expert Site Reliability Engineer
🇸🇦 Riyadh, Saudi Arabia🏢 On-site
Site Reliability EngineeringSREKubernetesTerraformCI/CDCloud PlatformsObservability
At a Glance
- Category
- 💻 Technology
- Level
- Mid-Level
- Type
- Full-time
Quick CV Check
- Get your ATS score for TAWANTECH in 30 seconds
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
Purpose:
To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.
Main Duties and Responsibilities:
- Define and implement advanced reliability engineering practices across critical technology services.
- Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
- Design automation to reduce manual operational activities and improve system resilience.
- Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
- Lead technical analysis and resolution of complex production incidents.
- Conduct root-cause analysis and drive permanent corrective and preventive actions.
- Design solutions to improve system availability, scalability, capacity, and disaster resilience.
- Identify reliability risks and recommend architectural and engineering improvements.
- Drive performance engineering and capacity planning for critical services.
- Provide advanced technical guidance and mentorship on SRE practices.
- Promote automation and engineering approaches that reduce operational toil and improve service reliability.
QUALIFICATIONS & REQUIREMENTS
- Bachelor’s degree in Computer Science, Software Engineering, IT, or a related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
- Strong experience in cloud platforms, Kubernetes, and production environments.
- Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
- Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
- Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
- Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
- Experience driving reliability improvements and reducing operational toil through automation.
- Strong analytical, problem-solving, and technical leadership skills.
- Experience in Banking, FinTech, or Payment environments is preferred.
Requirements
- •Bachelor’s degree in Computer Science, Software Engineering, IT, or a related field
- •5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles
- •Strong experience in cloud platforms, Kubernetes, and production environments
- •Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics
- •Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform)
- •Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA)
- •Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery
- •Strong analytical, problem-solving, and technical leadership skills
Nice to Have
- •Experience in Banking, FinTech, or Payment environments
Responsibilities
- •Define and implement advanced reliability engineering practices across critical technology services
- •Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets
- •Design automation to reduce manual operational activities and improve system resilience
- •Develop and enhance monitoring, observability, alerting, and incident detection capabilities
- •Lead technical analysis and resolution of complex production incidents
- •Conduct root-cause analysis and drive permanent corrective and preventive actions
- •Design solutions to improve system availability, scalability, capacity, and disaster resilience
- •Identify reliability risks and recommend architectural and engineering improvements
Related Jobs4 similar jobs
Browse Similar
Apply Now
Quick CV Check
- Get your ATS score for TAWANTECH in 30 seconds
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
GCC Info
Company
TAWANTECH offers technology solutions aimed at driving digital transformation for businesses. They provide services to help organizations innovate and grow.
Apply Now



