Introduction
We are seeking an AI Factory Operations Design and Management Lead to establish and coordinate the operating model for a large-scale, multi-tenant AI Data Center and GPU service environment. The role combines operations governance, service-management design, stakeholder coordination, incident and issue management, communications, program execution, and continuous improvement.
This position will work closely with infrastructure architects, platform engineers, security teams, service desk / NOC / SOC functions, business stakeholders, customers, and technology vendors. The lead is expected to translate operational requirements into clear processes, ownership models, service reviews, readiness criteria, communications, and management reporting while partnering with technical specialists for detailed platform implementation.
The successful candidate should be comfortable leading complex cross-functional initiatives, facilitating executive and operational discussions, managing high-visibility issues, and driving structured execution in a technology environment. Deep hands-on engineering expertise across every AI infrastructure platform is not required, but sufficient technical literacy to understand dependencies, risks, service impacts, and escalation requirements is important.
Your Role And Responsibilities
Operating Model, Governance, and Service Management
- Lead the definition and ongoing improvement of the AI Factory operating model across Service Management, Service Desk, NOC, SOC, Platform Operations, Network Operations, Storage Operations, SRE, customer teams, and vendors.
- Facilitate ownership, RACI, L1/L2/L3 support boundaries, escalation paths, governance forums, and service-review mechanisms with technical and business stakeholders.
- Coordinate the design and adoption of Incident, Major Incident, Problem, Change, Request, Release, Defect, and Knowledge Management processes with relevant process owners and technical leads.
- Define clear communications, decision, approval, maintenance-window, escalation, and closure requirements for operational workflows.
- Create governance routines that provide management visibility into service health, risks, actions, dependencies, decisions, and continuous-improvement priorities.
Stakeholder, Incident, and Service Communications
- Lead customer-facing operational workshops, service reviews, readiness meetings, and cross-functional governance sessions.
- Coordinate communications for major incidents and high-priority operational issues, including stakeholder updates, executive summaries, action tracking, and post-incident follow-up.
- Build clear communication protocols for outages, degraded service, planned changes, maintenance activities, security events, and service restoration.
- Ensure technical findings, service risks, decisions, and action items are translated into concise, audience-appropriate communications for executives, customers, operations teams, and vendors.
- Support issue and crisis-management activities by maintaining accurate status, ownership, timelines, decision logs, and escalation communications.
SOC, NOC, Monitoring, and Operational Coordination
- Work with SOC, NOC, SRE, security, platform, network, and storage specialists to define alert intake, triage, escalation, notification, investigation, and closure workflows.
- Coordinate monitoring and service-health requirements for GPU, servers, networks, storage, platform services, and customer-facing service components.
- Help define alert severity models, operational dashboards, service-reporting formats, ticketing requirements, and management reporting.
- Coordinate operational readiness for monitoring integration, alert routing, automated ticket creation, on-call responsibilities, and vendor escalation.
- Ensure operational procedures include appropriate audit evidence, communication checkpoints, and stakeholder notification requirements.
Platform Operations Governance
- Partner with platform and infrastructure specialists to define operational procedures, ownership, escalation criteria, and service controls for NVIDIA GPU infrastructure and related AI platform services.
- Coordinate governance requirements for cluster onboarding, tenant setup, resource allocation, network provisioning, storage access, service changes, recovery, and decommissioning.
- Maintain an operational dependency view across compute, GPU, network, storage, security, platform software, monitoring, and vendor support.
- Ensure vendor-escalation paths, support contacts, required diagnostic evidence, and service responsibilities are documented and reviewed.
- Support platform-specific runbooks for technologies such as NVIDIA, Rafay, Netris, Weka, or comparable products in collaboration with subject-matter experts.
SLA, KPI, Readiness, and Continuous Improvement
- Coordinate definition and reporting of service KPIs and operational indicators such as availability, incident response, restoration time, provisioning success, change success, backlog, and service-request performance.
- Facilitate service scorecards, operational reviews, risk reviews, RCA follow-up, known-error tracking, and continuous-improvement backlogs.
- Develop and govern Operational Readiness Review, service-transition, launch, handover, and acceptance checklists.
- Coordinate knowledge transfer, operational rehearsals, tabletop exercises, failure simulations, and recovery exercises with technical owners.
- Use structured reporting and stakeholder feedback to identify process gaps and drive measurable improvements in operational effectiveness.
Documentation and Program Execution
- Develop and maintain operating-model documents, governance charters, RACI matrices, escalation matrices, service-review packs, SOP frameworks, runbook standards, readiness checklists, and transition plans.
- Plan and track cross-functional workstreams, milestones, dependencies, risks, decisions, and action items across customer, internal, and vendor teams.
- Provide structured status reporting to management and customers, with clear articulation of outcomes, unresolved risks, decisions required, and next actions.
- Facilitate alignment across multicultural and multi-stakeholder teams and ensure agreed actions are followed through to closure.
- Support enablement, onboarding, knowledge-sharing, and operational communication programs required for service adoption and Day-2 readiness.
Preferred Education
Master's Degree
Required Technical And Professional Expertise
Degree in Information Technology, Computer Science, Engineering, Business, Communications, Service Management, or a related field, or equivalent professional experience.
- 10+ years of experience in a global or enterprise environment, with substantial responsibility for program management, operations governance, service management, stakeholder management, issue management, communications, or comparable cross-functional leadership.
- Demonstrated ability to design and execute strategies, programs, governance mechanisms, communications, and operating routines across complex organizations.
- Strong project / program management capability, including planning, prioritization, multi-workstream coordination, risk and issue tracking, and delivery under time-sensitive conditions.
- Experience working with senior management and multiple stakeholder groups, with the ability to provide consultative advice, facilitate decisions, and drive alignment.
- Experience handling high-visibility issues, incident or crisis communications, escalations, or other time-critical situations requiring structured coordination and clear messaging.
- Ability to create clear management reports, operational documentation, process materials, meeting outputs, and stakeholder communications.
- Ability to work effectively in a technology-driven environment and develop sufficient working knowledge of infrastructure, cloud, AI, security, network, and service-management concepts.
- Full professional proficiency in Chinese for customer workshops, operational coordination, documentation, service reviews, and issue communications.
- Working English proficiency for global collaboration, technical documentation, management reporting, and vendor communications.
- Ability to work locally in Taiwan and provide regular on-site support at customer offices and AI Data Center facilities.
- Availability for scheduled maintenance windows and critical-incident coordination when required.
Preferred Technical And Professional Experience
Experience in IT operations, managed services, ITSM, SRE, NOC, SOC, service delivery, or infrastructure operations governance.
- Working knowledge of Incident, Problem, Change, Request, Release, Knowledge Management, major-incident coordination, and service-transition practices.
- Experience defining or managing SLA, SLO, SLI, MTTA, MTTR, RTO, RPO, service scorecards, or operational KPI reporting.
- Exposure to AI infrastructure, GPU platforms, cloud platforms, data centers, multi-tenant services, or 24x7 technology environments.
- Familiarity with NVIDIA GPU infrastructure and platforms such as Rafay, Netris, Weka, Kubernetes, or comparable technologies; deep engineering expertise is not mandatory for this role.
- Experience developing SOPs, runbooks, readiness criteria, operational handover plans, service catalogs, or escalation frameworks.
- Relevant certifications in ITSM / ITIL, SRE, project or program management, security, audit, cloud, or infrastructure operations.