Site Reliability Engineer
tao digital solutions- Posted 2 hours ago
- Be among the first 10 applicants
Job Description
SRE Engineer - Multi Cloud
Own diagnosis and resolution across the multicloud estate: AWS, GCP, AliCloud, and the internal cloud gateway and control planes. You will take incidents from page or ticket intake through to resolution and written RCA, escalating to platform engineering only when a confirmed defect or capacity change is required. Singapore is the anchor site for this role, covering the APAC estate at its point of origin.
Cloud platform core — must have
• Account lifecycle across providers: onboarding, org hierarchy placement, service enablement, decommission
• Provisioning failure diagnosis and stuck-state recovery, from request intake to delivered resource
• Quota management: request, validation, pre-requisite checks, and capacity escalation to the provider
• IAM and access: roles, trust policies, federated and short-lived credentials, SSO role mapping
• Cost visibility and remediation: identifying waste, right-sizing, and tracking realized savings
Multi-cloud — must have at least two
• AWS: EC2, VPC, NLB and ALB, Route 53, subnet and address-space reconciliation, IAM roles anywhere
• GCP: GCE, service accounts, vulnerability remediation workflows, project and folder structure
• AliCloud: region-specific service availability and version constraints, STS and RAM
• Cross-provider differences in quota, IAM, and network models. Knowing where the three diverge is the point of the role
Platform tooling — must have
• Infrastructure as code: Terraform or Ansible, including drift detection and blast-radius management
• Deployment and release tooling such as Spinnaker
• Compass compliance onboarding, backup services, image and golden-AMI lifecycle
• Internal cloud inventory and query tooling for account, resource, and spend reporting
Observability — must have
• Prometheus and Grafana: dashboard authorship and alert definition that is actionable, not noise
• Splunk: forwarding, query authorship, and use in live incident diagnosis
• Defining SLIs for provisioning success, gateway availability, and request latency
Networking — must have
• Cross-environment DNS resolution and connectivity to managed database services
• Load balancer reconcile loops and subnet annotation behavior
• Gateway and proxy endpoint troubleshooting, including destination allow-listing and egress paths
• Outage triage and formal RCA authorship
Automation and code — must have
• Python or Go sufficient to ship tooling that goes through peer review, not throwaway scripts
• Converting repeat manual remediation into runbook automation or self-service paths
• Git-based operational workflow, GitHub Actions, PagerDuty
Incident discipline — must have
• Run an incident end to end: triage, mitigate, verify recovery against the customer-facing signal, resolve
• Journal every investigation in writing as you work, so the next shift continues instead of restarting
• Blameless RCA authorship and promotion of findings into the shared runbook library
• Clear written English. Most hand-offs happen in writing across time zones
Required Qualifications
• 5+ years in cloud infrastructure support, site reliability, or cloud operations engineering. The multi-provider scope sets this bar above a single-cloud equivalent role.
• Deep production expertise in at least one major cloud provider, with demonstrated working competence in a second. Depth in AWS or GCP is the most common qualifying profile.
• Strong identity and access management fundamentals that transfer across providers: role assumption and delegation, service and machine identities, policy evaluation and troubleshooting, and federated or directory-sourced group mapping.
• Practical Kubernetes operations experience (EKS, GKE, or ACK) — inspecting pod and node state, diagnosing volume and networking failures, and interpreting cluster events.
• Working knowledge of enterprise network fundamentals: routing, DNS, TLS, proxies, CIDR planning, and firewall/ACL models.
• Scripting proficiency in Python or Bash, plus fluency with provider CLIs.
• Excellent written English — the majority of support happens asynchronously in text, and clarity directly determines resolution speed.
• Legally authorized to work in Singapore.
Preferred Qualifications
• Certification in one or more of: AWS Solutions Architect (Associate/Professional), Google Professional Cloud Architect, Alibaba Cloud ACP.
• Prior AliCloud experience, or demonstrated rapid ramp on an unfamiliar cloud provider. AliCloud talent is scarce in the market; evidence of learning a new provider quickly is an acceptable and expected substitute.
• Experience supporting internal developer platforms in a large enterprise.
• Familiarity with hybrid connectivity between corporate networks and public cloud.
• Infrastructure-as-code exposure (Terraform, CloudFormation).
• Observability tooling experience (Splunk, Datadog, CloudWatch, Cloud Logging).
• Prior follow-the-sun or multi-region support rotation experience.
• Familiarity with China-region cloud operations and their regulatory constraints.

