AWS Platform Engineer (MF)
DVT
Pretoria, Gauteng
Job description
DVT is one of the top software development companies on the continent. Our software engineers are consulting on cutting edge applications at top companies in South Africa, as well as consulting globally. You will have the opportunity to work alongside some of the most established developers in the country and globally with the latest technologies.
DVT is committed to continuously training our staff and we are very proud of our culture of learning, internal speaking and training at a variety of sponsored technical events across the AWS ecosystem.
We are looking for a AWS Platform Engineer to join our cloud team.
You will work closely with cross-functional teams to ensure the smooth integration and deployment of applications, improve efficiency through automation, and implement best practices for continuous integration and delivery.
This is a client-facing consulting role where you will engage directly with enterprise clients across financial services, telecommunications, government, and other sectors. You will provide technical leadership, mentor junior engineers, and drive the adoption of best practices.
The ideal candidate is a problem solver with a strong technical background, excellent communication skills, and a passion for driving innovation in cloud infrastructure.
Role Summary
Own the production Kubernetes platform end-to-end; reliability, scalability, observability, and developer experience. Create and maintain the Azure DevOps build and deployment pipelines. This is a hands-on engineering role actively developing and executing infrastructure as code scripts working in development, user testing, and production environments.
DUTIES AND RESPONSIBILITIES
Platform Ownership
- Manage and upgrade production K8s clusters; enforce resource quotas, RBAC, and namespace policies
- Own cluster autoscaling, node pool configuration, and capacity planning
- Manage workloads and cluster lifecycle via Rancher; maintain multi-cluster visibility and access control
- Own Helm chart library — versioning, releases, and values management across environments
Reliability & Uptime
- Define and track SLOs/SLIs for critical services
- Own on-call rotation; lead incident response and post-mortems
- Proactively identify and eliminate single points of failure
Observability
- Maintain and improve the monitoring stack (Prometheus, Grafana, Loki or equivalent)
- Build and tune alerting rules; reduce alert noise•Ensure distributed tracing is in place for key services
Pipeline & Delivery
- Maintain CI/CD pipelines in Azure DevOps; enforce security scanning and policy gates
- Partner with dev teams to unblock delivery without compromising stability
Cost & Efficiency
- Right-size workloads; identify and eliminate waste
- Report on infrastructure spend with actionable optimisation recommendations
Required Experience and Skills
Must-have
- 4+ years in infrastructure/platform/SRE roles
- Production-grade K8s experience (CKA preferred or equivalent depth)
- Strong Terraform; remote state, modules, CI-driven apply on Azure
- Proficiency in at least one scripting/systems language: Go, Python, or Bash
- Hands-on incident response experience — you've been on-call, you've done post-mortems
- Solid networking fundamentals: TCP/IP, routing, DNS, TLS, ingress controllers, service mesh basics
- Linux administration fluency
Advantageous
- Rancher (multi-cluster management, RBAC, catalog apps)
- Helm chart authoring and release management
- Autoscaler
- Azure (AKS, networking, Key Vault, storage)
- Secrets management: Azure Key Vault or Vault
- Experience setting up or maturing an observability stack from scratch
- Windows server experience
Who we are:
Good to know
How do I apply for this job?
Tap "Apply on Indeed" to open the original listing, where you can read the full description and apply directly. JobsZA never charges you to apply, and you should never pay money to get a job.
Found on Indeed · Posted Today