Role Overview
We are looking for an experienced Senior Software Engineer or Technical Lead to design and build scalable machine learning platforms for training, optimizing, evaluating, packaging, and deploying large-scale AI models.
The role requires strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering. You will work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.
Key Responsibilities
- Design and develop distributed ML platform services and reusable libraries.
- Build scalable training capabilities for large language and multimodal models.
- Support data, tensor, pipeline, and model parallelism across multi-node GPU clusters.
- Improve training throughput, GPU utilization, memory efficiency, communication performance, and fault recovery.
- Define stable APIs and platform interfaces for model onboarding, training, evaluation, and deployment.
- Integrate model optimization techniques such as quantization, pruning, knowledge distillation, and compression.
- Build evaluation and artifact-management workflows to measure model quality and system performance.
- Develop automated validation, CI/CD, regression testing, observability, and release pipelines.
- Profile GPU workloads and resolve end-to-end system performance bottlenecks.
- Build production-operational mechanisms, including metrics, alarms, dashboards, runbooks, and root-cause corrective actions.
- Collaborate with model, compiler, runtime, hardware, security, infrastructure, and product teams.
- Prepare technical designs, evaluate architecture trade-offs, and drive engineering decisions.
- Mentor engineers and improve code reviews, design reviews, testing, and development practices.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- 5+ years of professional software development experience.
- Strong programming skills in C++, Java, C#, Python, or a similar language.
- Experience designing distributed systems, high-performance computing platforms, or scalable backend systems.
- Strong understanding of system design, concurrency, reliability, scalability, and performance optimization.
- Experience as a technical lead, architect, mentor, or engineering team lead.
- Experience across the software development lifecycle, including coding, testing, source control, build systems, deployment, and production support.
Preferred Qualifications
- Experience with PyTorch, TensorFlow, JAX, NeMo, or Megatron-LM.
- Experience building distributed ML training, inference, evaluation, or data platforms.
- Knowledge of CUDA, GPU kernels, GPU profiling, and performance optimization.
- Experience with containers, Kubernetes, cloud infrastructure, CI/CD, and observability tools.
- Knowledge of model compression, quantization, pruning, distillation, compilation, or edge AI deployment.
- Experience with large language models, multimodal models, and distributed GPU training.
- Experience designing extensible platform APIs and reusable software frameworks.
- Experience working with applied science, hardware, compiler, runtime, or product engineering teams.
Key Skills
Distributed ML Systems | GPU Computing | ML Infrastructure | PyTorch | CUDA | Kubernetes | CI/CD | Model Optimization | Performance Engineering | System Architecture | Technical Leadership