Skip to content

AI Engineer

  • On-site
    • Riyadh, Riyadh Province, Saudi Arabia
  • AZM ICT - PMO

Lead Applied AI Engineer to build production LLM, RAG & multi-agent systems while owning cloud, MLOps, CI/CD, Kubernetes, security, observability, and technical leadership.

Job description

AI systems

  • Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).

  • Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95).

  • Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.

  • Implement safety guardrails — input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions.

  • Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders.

Platform & infrastructure

  • Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes.

  • Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments.

  • Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets.

  • Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines.

  • Own disaster recovery design — automated backups, failover, and documented RTO/RPO commitments.

Engineering leadership

  • Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets.

  • Standardize SDLC practices — branching strategy, PR governance, release management, delivery reporting — to improve lead time and deployment frequency.

  • Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery.

  • Produce handover documentation and runbooks that make systems auditable and operationally transferable.

  • Support vendor and licensing negotiations for cloud enterprise agreements.

Job requirements

Required Qualifications

  • Bachelor's degree in Software Engineering, Computer Science, or a related field.

  • 6–8+ years in software, DevOps, or platform engineering, including at least 2 years in an applied AI or ML engineering capacity.

  • Proven delivery of production AI/LLM systems — not only research or notebook-stage work.

  • Strong Python; comfortable with Bash and YAML.

  • Deep hands-on experience with Kubernetes, Docker/Podman, and Terraform.

  • Production experience with at least one major cloud (Azure preferred; OCI or GCP acceptable).

  • Demonstrated ownership of CI/CD at scale (Azure DevOps, GitHub Actions) and GitOps release models.

  • Experience leading a team and setting engineering standards across multiple squads.

Preferred Qualifications

  • Master's degree in Applied AI, Machine Learning, or a related discipline.

  • Fine-tuning experience with QLoRA/LoRA on GPU clusters; PyTorch and Transformers.

  • Vector database experience (Milvus, Pinecone, or Weaviate) and RAG retrieval design.

  • Experience delivering on Saudi government or large-scale national digital platforms, with familiarity in local compliance and standards.

  • Arabic and English professional proficiency.


Technical Environment

Python, FastAPI, PyTorch, Transformers, LangGraph, Milvus/Pinecone/Weaviate, Redis, PostgreSQL, Kubernetes, Docker/Podman, Terraform, Argo CD, Azure DevOps, GitHub Actions, Azure ML, Azure Monitor / Application Insights, ELK, SonarQube, Black Duck, Fortinet FW/WAF.

or