Search by job, company or skills

gmi cloud

Infra Engineer - SRE

Save
  • Posted 13 hours ago
  • Be among the first 10 applicants
Early Applicant

Job Description

About GMI

GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only six cloud providers worldwide to earn NVIDIA's prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to build AI without limits, providing everything they need to prototype, train, and deploy AI models quickly and reliably.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  1. Design, implement and maintain scalable AI/ML infrastructure solutions.
  2. Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  3. Automate deployment, configuration and management of infrastructure resources.
  4. Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  5. Implement CI/CD pipelines for infrastructure deployment and orchestration.
  6. Ensure security, compliance and best practices across infrastructure.
  7. Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  8. Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  9. Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  10. Regional/international travel to GMI data center locations.

Qualifications

  1. Bachelor's degree in Computer Science or related field.
  2. Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  3. Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  4. Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  5. Familiarity with Linux system administration and scripting (Python, Bash).
  6. Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  7. Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  8. Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  9. Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you're excited about this role, we'd love to hear from you. We believe diverse perspectives and experiences strengthen our team.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151473087