DevOps / Platform Engineer

Habib Qayoom

What I do

With 2+ years hands-on in cloud infrastructure, I run production Kubernetes on Azure: end-to-end observability, incident response, and infrastructure as code, plus a monitoring bill I cut by 71%.

4
Production AKS clusters
40+
Grafana dashboards owned
71%
Monitoring spend cut

Experience

Where I've built

Three roles, widening scope: from first Dockerfile to owning production.

Certification

AWS Certified Solutions Architect - Associate (SAA-C03)

Amazon Web Services · ActiveVerify on Credly

Techanzy (contracted to Implement AI, AIOS Platform, London)

Aug 2025 to Present

DevOps / Platform Engineer

  • Operate production infrastructure for an enterprise AI SaaS platform across 4 Azure AKS clusters, PostgreSQL, Azure Functions, Key Vault, and Blob Storage.
  • Engineered a secure, isolated Kubernetes workload platform for browser automation: namespace-scoped ServiceAccount authentication, dedicated node-pool isolation via taints and tolerations and node selectors, and automated provisioning and deletion of session Jobs and Ingresses with cert-manager TLS.
  • Integrated the AIOS backend with the AKS API using ServiceAccount token authentication and CA certificate validation instead of admin kubeconfig access, with Kubernetes Secrets for secure credential management.
  • Led production incident response: diagnosed and resolved FastAPI/Gunicorn OOMKills, high memory usage, and production latency issues across services.
  • Built and maintain 40+ Grafana Cloud dashboards using Prometheus, Loki, Tempo, and Application Insights for production logs, metrics, LLM usage, calls, and billing.
  • Built a Python-based Loki to Slack alerting pipeline with deduplication, suppression windows, and a manual-review web UI; created 30+ Grafana alert rules routed to Slack and email.
  • Cut Grafana Cloud spend by roughly 71% (from about $1,090/mo overage to about $316/mo) by analyzing root causes of metric growth and applying targeted optimization strategies.
  • Self-hosted and operated the Trieve RAG platform on AKS, including Qdrant, embedding and reranker models, and Keycloak SSO, alongside browser-use AI agents and n8n automation.
  • Drove ISO 27001:2022 compliance evidence and remediation activities using Vanta.

Code to Kloud

Jan 2025 to Aug 2025

DevOps Engineer

  • Provisioned AWS infrastructure with Terraform (VPC, multi-AZ subnets, NAT/IGW, security groups) and EKS clusters (managed node groups, ALB Controller via Helm with IAM OIDC/IRSA, EBS CSI plus PV/PVC); enforced namespace-scoped RBAC via the aws-auth ConfigMap.
  • Built a CI/CD pipeline for ECS Fargate (ECR SHA-tagged build/push, AWS Secrets Manager, task-definition re-registration, stability gating); centralized logs to Loki via FireLens and deployed EKS workloads with ArgoCD.
  • Implemented multi-environment promotion (dev, staging, prod) with branch-based workflows and environment-scoped approvals, keeping deployments consistent and auditable across environments.

Upvave

Jun 2024 to Dec 2024

DevOps Intern → Associate DevOps Engineer

  • Architected AWS networking (VPC, multi-AZ subnets, NAT/IGW, security groups) provisioned with Terraform; deployed workloads on EKS and ECS Fargate with GitHub Actions and AWS CodePipeline, including CodeDeploy blue/green pipelines on ECS.
  • Deployed an observability stack (Prometheus, Grafana, Loki, Fluent Bit) across EKS/ECS clusters; managed secrets with AWS Secrets Manager and IAM; implemented event-driven automation with AWS Lambda.
  • Containerized applications with Docker and automated image build and push to ECR through GitHub Actions, standardizing deployment artifacts across services.

Selected work

Projects & case studies

Grouped by type. Each one opens a full breakdown: the problem, what I did, and the result.

Production systems I run today.

Featured

AIOS Platform

Production multi-client AI platform on Azure AKS.

AzureAKSPostgreSQLGrafana Cloud
Case studyLive

Also running

More of the Implement AI platform

Internal tools that support AIOS in production. Not publicly reachable, so there's nothing to link to here, just what they are.

Admin App

Internal

Internal admin console for the Implement AI platform.

Core

Internal

Internal management tool for the Implement AI platform.

Computer Use Agent

Internal

Internal computer-use AI agent tooling.

Charis

Internal

A custom application built for a specific client.

n8n Automations

Internal

Workflow automation service powering platform integrations and pipelines.

Trieve RAG

Internal

Self-hosted retrieval-augmented generation stack on AKS.

Toolkit

Tech stack

The cloud platforms, observability tools, and infrastructure-as-code I use to keep production systems reliable and cost-efficient.

Cloud & Infrastructure
AzureAKSAzure FunctionsKey VaultPostgreSQL Flexible ServerBlob StorageEntra IDAWS EC2AWS EKSAWS ECSVPCIAMS3LambdaCloudflare Workers
Containers & Orchestration
KubernetesDockerHelmcert-managerACRHPA / Cluster Autoscaler
Observability & Monitoring
Grafana CloudPrometheus / PromQLLoki / LogQLTempoPyroscopeGrafana AlloyApplication InsightsSynthetic Monitoring
IaC & CI/CD
TerraformGitHub ActionsAWS CodePipeline / CodeDeployJenkinsArgoCDAzure DevOpsDependabot
Scripting & Databases
PythonBashPowerShellPostgreSQLMySQLSQLKQL
Security & Compliance
ISO 27001:2022 (Vanta)CIS BenchmarksNIST 800-53OWASP ZAPBanditRBACSecrets management

Let's talk

Have infrastructure that needs owning?

I'm open to DevOps, SRE, and platform roles, full-time or contract. Based in Lahore (UTC+5) and used to working across UK and US time zones. Email gets the fastest response.

Download resume