Search Jobs

Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

tao digital solutions
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

SRE Engineer - Multi Cloud

Own diagnosis and resolution across the multicloud estate: AWS, GCP, AliCloud, and the internal cloud gateway and control planes. You will take incidents from page or ticket intake through to resolution and written RCA, escalating to platform engineering only when a confirmed defect or capacity change is required. Singapore is the anchor site for this role, covering the APAC estate at its point of origin.

Cloud platform core — must have

• Account lifecycle across providers: onboarding, org hierarchy placement, service enablement, decommission

• Provisioning failure diagnosis and stuck-state recovery, from request intake to delivered resource

• Quota management: request, validation, pre-requisite checks, and capacity escalation to the provider

• IAM and access: roles, trust policies, federated and short-lived credentials, SSO role mapping

• Cost visibility and remediation: identifying waste, right-sizing, and tracking realized savings

Multi-cloud — must have at least two

• AWS: EC2, VPC, NLB and ALB, Route 53, subnet and address-space reconciliation, IAM roles anywhere

• GCP: GCE, service accounts, vulnerability remediation workflows, project and folder structure

• AliCloud: region-specific service availability and version constraints, STS and RAM

• Cross-provider differences in quota, IAM, and network models. Knowing where the three diverge is the point of the role

Platform tooling — must have

• Infrastructure as code: Terraform or Ansible, including drift detection and blast-radius management

• Deployment and release tooling such as Spinnaker

• Compass compliance onboarding, backup services, image and golden-AMI lifecycle

• Internal cloud inventory and query tooling for account, resource, and spend reporting

Observability — must have

• Prometheus and Grafana: dashboard authorship and alert definition that is actionable, not noise

• Splunk: forwarding, query authorship, and use in live incident diagnosis

• Defining SLIs for provisioning success, gateway availability, and request latency

Networking — must have

• Cross-environment DNS resolution and connectivity to managed database services

• Load balancer reconcile loops and subnet annotation behavior

• Gateway and proxy endpoint troubleshooting, including destination allow-listing and egress paths

• Outage triage and formal RCA authorship

Automation and code — must have

• Python or Go sufficient to ship tooling that goes through peer review, not throwaway scripts

• Converting repeat manual remediation into runbook automation or self-service paths

• Git-based operational workflow, GitHub Actions, PagerDuty

Incident discipline — must have

• Run an incident end to end: triage, mitigate, verify recovery against the customer-facing signal, resolve

• Journal every investigation in writing as you work, so the next shift continues instead of restarting

• Blameless RCA authorship and promotion of findings into the shared runbook library

• Clear written English. Most hand-offs happen in writing across time zones

Required Qualifications

• 5+ years in cloud infrastructure support, site reliability, or cloud operations engineering. The multi-provider scope sets this bar above a single-cloud equivalent role.

• Deep production expertise in at least one major cloud provider, with demonstrated working competence in a second. Depth in AWS or GCP is the most common qualifying profile.

• Strong identity and access management fundamentals that transfer across providers: role assumption and delegation, service and machine identities, policy evaluation and troubleshooting, and federated or directory-sourced group mapping.

• Practical Kubernetes operations experience (EKS, GKE, or ACK) — inspecting pod and node state, diagnosing volume and networking failures, and interpreting cluster events.

• Working knowledge of enterprise network fundamentals: routing, DNS, TLS, proxies, CIDR planning, and firewall/ACL models.

• Scripting proficiency in Python or Bash, plus fluency with provider CLIs.

• Excellent written English — the majority of support happens asynchronously in text, and clarity directly determines resolution speed.

• Legally authorized to work in Singapore.

Preferred Qualifications

• Certification in one or more of: AWS Solutions Architect (Associate/Professional), Google Professional Cloud Architect, Alibaba Cloud ACP.

• Prior AliCloud experience, or demonstrated rapid ramp on an unfamiliar cloud provider. AliCloud talent is scarce in the market; evidence of learning a new provider quickly is an acceptable and expected substitute.

• Experience supporting internal developer platforms in a large enterprise.

• Familiarity with hybrid connectivity between corporate networks and public cloud.

• Infrastructure-as-code exposure (Terraform, CloudFormation).

• Observability tooling experience (Splunk, Datadog, CloudWatch, Cloud Logging).

• Prior follow-the-sun or multi-region support rotation experience.

• Familiarity with China-region cloud operations and their regulatory constraints.

More Info

Key Skills

AliCloud

PagerDuty

GitHub Actions

Spinnaker

Similar Jobs

5-7 yrs
Taiwan, Taipei City
Skills:
ElkPrometheusBashGrafanaDevopsVMWareTerraformAnsiblePythonAWSEFKOpenTelemetryInfrastructure EngineeringLinux systems
5-7 yrs
Remote
Skills:
Incident ManagementTest CasesDistributed SystemsPythonAccess ManagementCapacity Planning