Search Jobs

Search by job, company or skills

Server Engineer AI Cluster/GPU

Server Engineer AI Cluster/GPU

palebluedot ai
5-7 Years
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities

  • Manage the full lifecycle of data center servers, including design, deployment, performance tuning, and validation, as well as the implementation, operation, and maintenance of HPC/AI clusters.
  • Lead the deployment of heterogeneous computing resources such as GPUs and XPUs, optimize system performance and stability, and monitor and maintain server operations.
  • Prepare technical documentation, provide customer support, and drive the development of automation scripts using Shell, Python, and Ansible.

Key Requirements

  • At least 5 years of relevant experience, with hands-on expertise in building large-scale HPC clusters with thousands of GPUs, GPU hardware architecture, and parallel computing technologies such as MPI and OpenMP.
  • Strong proficiency in virtualization and containerization technologies, with experience in HPL and NCCL benchmarking and high-performance file systems such as Lustre and GPFS.
  • HPC-related certifications are preferred.
  • Professional working proficiency in English.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

containerization technologies

HPC clusters

HPL

GPUs

GPU hardware architecture

NCCL

parallel computing technologies

Shell

About Company