
AI Engineer
At a Glance
- Category
- 💻 Technology
- Level
- Mid-Level
- Type
- Full-time
Stop — Check First
- Check your resume before InnovationTeam rejects it
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
Looking for an AI Engineer to Develop, deploy, and operate AI/LLM models across Clinets dual environment — GCP for public-cloud workloads, Humain sovereign cloud for classified data.
Build and fine-tune LLM/ML models for Arabic NLP, document classification, vision/OCR, and AIOps use cases.
Run pre-deployment evaluation
Accuracy baselines, regression and safety testing; evidence to justify GPU allocation.
Optimize inference — quantization, batching, context sizing — against measured usage.
Deploy on Humain GPUaaS: Kubernetes, GPU partitioning on B300 nodes, quotas, RBAC.
Build equivalent workloads on GCP (Vertex AI, GKE) with classification-based routing.
Own serving stack (vLLM/TGI), model versioning, CI/CD, and monitoring for latency, tokens, GPU utilization, and drift.
Ensuring developed AI Models Complying with ZATCA data sovereignty and SDAIA requirements (AI Ethics, GenAI Guidelines, PDPL).
5 years ML/AI engineering, in production LLM deployment with knowledge in
Python, PyTorch, Hugging Face
Kubernetes in production; GPU-served inference
GCP Vertex AI or any equivellent cloud
Requirements
- •5 years of ML/AI engineering experience
- •Experience in production LLM deployment
- •Proficiency in Python
- •Knowledge of PyTorch
- •Hugging Face experience
- •Kubernetes in production
- •GPU-served inference experience
- •GCP Vertex AI or equivalent cloud experience
Responsibilities
- •Develop, deploy, and operate AI/LLM models across dual environments (GCP and Humain sovereign cloud)
- •Build and fine-tune LLM/ML models for Arabic NLP, document classification, vision/OCR, and AIOps
- •Run pre-deployment evaluation Accuracy baselines, regression, and safety testing
- •Optimize inference via quantization, batching, and context sizing
- •Deploy on Humain GPUaaS using Kubernetes and GPU partitioning on B300 nodes
- •Build equivalent workloads on GCP (Vertex AI, GKE) with classification-based routing
- •Own serving stack (vLLM/TGI), model versioning, CI/CD, and monitoring
- •Ensure AI Models comply with ZATCA data sovereignty and SDAIA requirements
Related Jobs4 similar jobs
Browse Similar
- Check your resume before InnovationTeam rejects it
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.


