
AI Engineer
- On-site
- Riyadh, Riyadh Province, Saudi Arabia
- AZM ICT - PMO
Lead AI-native platform design and delivery, covering LLMs, multi-agent systems, cloud infrastructure, CI/CD, GitOps, observability, security, and technical team leadership.
Job description
Required Qualifications:
Bachelor's degree in Software Engineering, Computer Science, or a related field.
6–8+ years in software, DevOps, or platform engineering, including at least 2 years in an applied AI or ML engineering capacity.
Proven delivery of production AI/LLM systems — not only research or notebook-stage work.
Strong Python; comfortable with Bash and YAML.
Deep hands-on experience with Kubernetes, Docker/Podman, and Terraform.
Production experience with at least one major cloud (Azure preferred; OCI or GCP acceptable).
Demonstrated ownership of CI/CD at scale (Azure DevOps, GitHub Actions) and GitOps release models.
Experience leading a team and setting engineering standards across multiple squads.
Preferred Qualifications:
Master's degree in Applied AI, Machine Learning, or a related discipline.
Fine-tuning experience with QLoRA/LoRA on GPU clusters; PyTorch and Transformers.
Vector database experience (Milvus, Pinecone, or Weaviate) and RAG retrieval design.
Experience delivering on Saudi government or large-scale national digital platforms, with familiarity in local compliance and standards.
Arabic and English professional proficiency
Job requirements
Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).
Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95).
Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.
Implement safety guardrails — input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions.
Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders.
Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes.
Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments.
Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets.
Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines.
Own disaster recovery design — automated backups, failover, and documented RTO/RPO commitments.
Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets.
Standardize SDLC practices — branching strategy, PR governance, release management, delivery reporting — to improve lead time and deployment frequency.
Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery.
Produce handover documentation and runbooks that make systems auditable and operationally transferable.
Support vendor and licensing negotiations for cloud enterprise agreements
or
All done!
Your application has been successfully submitted!
You've already applied for this job
We appreciate your interest in this position. Unfortunately, you have already applied for this job.
