

Search by job, company or skills
Position Overview
We are looking for an experienced Senior CloudOps Engineer to join our infrastructure team. You will be responsible for designing, operating, and continuously improving a highly available cloud platform that supports engineering teams across multiple regions and continents. This role requires deep technical expertise, a strong operational mindset, and the ability to collaborate closely with software engineers to deliver scalable, secure, and reliable infrastructure.
Technical Skills & Qualifications
• Strong expertise in networking fundamentals, including:
• Extensive hands-on experience administering and troubleshooting Linux systems in production environments.
• Strong experience with Amazon Web Services (AWS), including designing, deploying, and operating cloud-native infrastructure.
• Experience building, maintaining, and optimizing CI/CD pipelines for automated software delivery.
• Strong knowledge of Kubernetes, preferably Amazon EKS, including cluster administration, deployments, scaling, and troubleshooting.
• Experience using GitHub, including Git workflows, repository management, GitHub Actions, and collaborative development practices.
• Ability to automate operational tasks using scripting languages such as Bash or Python.
• Experience with infrastructure monitoring, alerting, and observability tools.
• Strong troubleshooting skills with the ability to quickly identify root causes across infrastructure, networking, Kubernetes, and application layers.
• Excellent understanding of infrastructure reliability, high availability, disaster recovery, and scalability principles.
• Strong documentation and communication skills.
Skills, that are a plus:
• Experience with Terraform and Infrastructure as Code (IaC).
• Understanding of cloud security best practices, including:
• Experience with container security and Kubernetes security best practices.
• Familiarity with observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, or similar tools.
• Experience supporting multi-account AWS environments.
• Experience with incident response, postmortems, and operational excellence practices.
• Experience leveraging AI tools, particularly the Claude API, to improve operational efficiency. This includes understanding how to write effective prompts, provide appropriate context, filter and validate responses, and use AI to accelerate troubleshooting, documentation, automation, and engineering workflows rather than treating it as a simple chatbot.
Key Responsibilities
What We're Looking For
Job ID: 152343547