Found Description
Responsibilities
- Lead end-to-end technical deployments for GPU neocloud and AI Factory customers
- Configure and troubleshoot bare metal GPU node infrastructure, including CNI, GPU Operator, and RDMA/InfiniBand
- Deploy and validate Kubernetes and vCluster environments
- Provide knowledge transfer to customer teams to ensure platform self-sufficiency
- Document reusable playbooks and deployment architectures
- Collaborate with Engineering and Product to surface infrastructure challenges
- Partner with Sales in the pre-sales process for deep infrastructure proof of value engagements
Requirements
- 5+ years of experience deploying and operating Kubernetes in production
- Practical knowledge of NVIDIA GPU Operators and CUDA tooling
- Deep understanding of CNI plugins, overlay networks, and load balancing
- Experience with persistent volume configuration, CSI drivers, and distribute...