Site Reliability Engineer (AI Factory)
Liquid Tech (Pty) Ltd.
Cape Town, Western Cape
Driving jobs span delivery, courier, code 10 and code 14 truck driving. A valid licence and PrDP are usually the only entry requirements, making this one of the most accessible skilled trades in South Africa.
This listing does not state a salary. As a guide, driver roles in South Africa typically pay R7 000 to R18 000 a month (indicative).
Job description
Role Purpose
Cassava Technologies is a pan-African digital services leader. We are building a state-of-the-art AI Factory in Cape Town, South Africa, based on the NVIDIA NCP (NVIDIA Cloud Partner) program. This facility will serve as the foundation for high-performance AI workloads across the continent. The Site Reliability Engineer (SRE) - Compute is responsible for engineering reliability into the Cassava AI Factory platform. The role focuses on reducing operational risk, improving service availability, and eliminating repetitive operational effort (toil) through automation, observability, and reliability-focused engineering. The SRE works closely with AI Cloud Operations Engineers and AI Engineering teams to ensure that the AI Factory becomes more stable, predictable, and scalable over time.You ensure the GPUs are healthy, updated, and never starved for data.
You will manage the extreme-throughput storage layers and the full lifecycle of the NVIDIA compute stack. Deep Expertise: High-Performance Storage (HPS) architectures (specifically Weka or similar) and deep knowledge of the NVIDIA GPU stack.AI Specifics: Managing NVIDIA drivers, firmware (SBIOS/VBIOS), and the CUDA software layer. Familiarity with local high-end NVMe disk pooling. Key Responsibilities: Executing hardware acceptance testing (fio, cluster validation), managing firmware/driver parity across the cluster, troubleshooting GPU health (NVLink errors, thermal throttling), and optimizing data ingress/egress for massive AI datasets.
What You Will Do Day-to-Day
? Monitor & Respond: Act as the first line of defense for the AI Factory, monitoring the health of the 2,048 GPU cluster and responding to complex, multi-layered incidents.
? Automate: Ruthlessly eliminate manual toil by building robust automation for provisioning, configuration management, and self-healing.
? Collaborate: Work seamlessly alongside hardware vendors (HPE, NVIDIA), the network engineering team, and software architects to bridge physical data center realities with cloud-native workflows.
? Scale: Assist in capacity planning, acceptance testing (UAT) for new hardware deliveries, and continuously tuning the environment for maximum throughput.
Role Description
Engineer platform reliability across AI Factory services (BMaaS, GPUaaS, LLM Training, AIFaaS, AISaaS).Define, measure, and improve Service Level Objectives (SLOs), and reliability targets. Identify systemic weaknesses that contribute to outages, performance degradation, or instability. Full Stack Oversight: Oversee the daily health of the environment, from the physical layer (Power/Cooling/Cabling) up through Compute/Storage (NVIDIA H200s/Infiniband Fabric, WEKA Storage) and the Orchestration Layer (Rafay).Act as a Tier 2 escalation point for complex incidents, working closely with the AI Centre of excellence engineering teams. Support Tier 1 Engineers during high severity incidents with deep technical analysis. Participate in on call rotation for critical escalations. Drive post incident reviews focused on learning and prevention. Work closely with vendors – ADC, WEKA, HPe, Nvidia, Rafay, LIT C2 for troubleshooting and incident resolution.
Identify repetitive, manual, or high effort operational activities that do not add long term value. Identify opportunities for automation to reduce recurring incidents and manual intervention. Improve tooling, scripts, self-healing mechanisms, and guardrails that reduce operational load. Design and maintain observability dashboards for availability and performance of the AI Factory. Improve monitoring, alerting, and dashboards to reduce noise and improve signal quality. Ensure alerts are actionable, severity aligned and mapped to operational response. Proactive Monitoring: Utilize observability tools to detect bottlenecks (thermal throttling, packet loss) before they disrupt customer training runs or inferencing requests. Collaborate with AI Engineering and vendors on platform design changes that improve reliability and recovery. Contribute to capacity modelling and scaling strategies. Identify reliability risks introduced by new features, changes, or capacity growth.
Operational Security: Enforce network security policies to ensure tenant isolation and data sovereignty. Maintenance Windows: Schedule and execute patching cycles (Firmware, NVIDIA Drivers, OS) with extreme care, ensuring coordination with customers to avoid killing active training jobs. Work closely with all levels of support tier’s to understand the incident behaviours and how to reduce repeat incidents. Compliance: Maintain strict operational adherence to NVIDIA NCP standards for availability and performance. Work closely with facility providers to ensure facility reliability is in place, ensure optimal uptime of cooling, power and security systems of the facility provider. Collaborate with Security teams on patching strategies and remediation workflows. Ensure secure configuration of infrastructure components (OS, Kubernetes, networking layers).Contribute to incident response for security-related events (breaches, compromise, data exposure)
Strong Linux systems engineering background, 7+ years of hands-on experience.
Minimum of 7 years experience operating distributed systems at scale including compute and storage.
Familiarity with container platforms (Kubernetes) and orchestration concepts, Hypervisor platforms.
Observability tooling (metrics, logging, tracing)
Automation and scripting skills (Python, Go, Bash, or similar)
Deep technical understanding of networking (InfiniBand and Ethernet) and storage concepts in high performance environments
Understanding of infrastructure security principles (network segmentation, zero trust, least privilege)
Bonus points for: NVIDIA NCP certifications, CKA/CKS, advanced networking certifications (CCNP/CCIE equivalent), or direct experience with Weka and Rafay.
Familiarity with identity and access management (IAM), RBAC, and authentication mechanisms
Knowledge of container and Kubernetes security (image scanning, runtime protection, policy enforcement)
Support Stack: Experience with ITSM/Ticketing tools (e.g., Jira Service Desk, ServiceNow)
Observability: Familiarity with monitoring tools to validate system health.
Deep technical experience of storage systems – SAN/NAS
Good to know
What does this driver job pay?
This listing does not state a salary. As a guide, driver roles in South Africa typically pay R7 000 to R18 000 a month (indicative).
Do I need experience for driver jobs in Cape Town?
This driver role may ask for some experience or a relevant qualification. Read the listing for the specifics before you apply.
How do I apply for this job?
Tap "Apply on Indeed" to open the original listing, where you can read the full description and apply directly. JobsZA never charges you to apply, and you should never pay money to get a job.
Found on Indeed · Posted Yesterday
More driver and similar jobs in Cape Town
Confidential
R28K - R28K/mo