At NVIDIA, we are pushing the boundaries of AI, graphics, and computing. The GitHub Actions Runner team manages self-hosted GPU-enabled GitHub Actions runners, using Actions Runner Controller with KubeVirt to deploy ephemeral VM-based runners for NVIDIA’s open source projects on GitHub. The team operates both on-premise and in the cloud (AWS) to support 100+ developers with whom they collaborate regularly to ensure a seamless CI/CD experience. We are looking for a Senior Infrastructure Engineer to help scale, optimize, and expand our platform.Our team is fully remote and distributed across multiple time zones. If you're passionate about infrastructure, Kubernetes, automation, and observability, this is an opportunity to work with exciting technology at one of the most innovative companies in the world.Preferred work location: Eastern/Central time zonesWhat you'll be doing:Manage and scale self-hosted GitHub Actions runners using KubernetesHelp expand runner support for various hardware and operating system combinations, including Linux, Windows, single-GPU, multi-GPU, NVLink, and moreUse Infrastructure as Code (Terraform and ArgoCD) to deploy and maintain infrastructure both on-premise and in AWSBuild and maintain runner VM images using HashiCorp PackerConnect distributed services securely using mTLS, PKI, and HashiCorp VaultDevelop, package, and deploy custom Golang tools to support platform observability, stability, and efficiencyConfigure alerting and monitoring to identify and address issues quickly, using tools like Prometheus and GrafanaContribute upstream to open-source tools and libraries that our team depends onPeriodically update platform dependencies and address CVEsWhat we need to see:B.S. or M.S. in Computer Science, Computer Engineering, or a related field (or equivalent experience)7+ years of proven experience in infrastructure, DevOps, or platform engineeringStrong Kubernetes expertise (running, debugging, and scaling workloads)Experience with GitOps tools (ArgoCD or similar)Proficiency in Linux administration and troubleshootingExperience with Infrastructure as Code using Terraform/TerragruntProficiency in Golang, Python, and TypeScriptHands-on experience with monitoring, logging, and tracing (Prometheus, Grafana, OpenTelemetry, etc.)Solid understanding of CI/CD pipelines, particularly GitHub ActionsAbility to work and collaborate effectively with a fully remote, distributed teamWays to stand out from the crowd:Experience instrumenting telemetry for distributed systemsStrong background in GPU workloads on Kubernetes with experience writing custom Kubernetes controllers Deep understanding of KubeVirt and/or virtualizationExperience with self-hosted GitHub Actions runnersContributions to open-source Kubernetes-related projectsWith competitive salaries and a generous benefits package, NVIDIA is considered one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the industry working for us. Due to unprecedented growth, our exclusive engineering teams are expanding rapidly. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!The base salary range is 168,000 USD - 333,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View Original Job Posting