Introduction
We are seeking an experienced AI Data Center Architect to lead solution architecture, integration planning, and technical governance for a large-scale, multi-tenant AI infrastructure platform. The role bridges business objectives, AI and data-platform requirements, infrastructure capabilities, and implementation teams to create an architecture that is scalable, secure, operable, and commercially practical.
The architect will translate customer requirements into High-Level Designs and coordinate Low-Level Designs with domain specialists covering GPU compute, networking, storage, Kubernetes, automation, security, monitoring, and Day-2 operations. The role requires strong customer-facing consulting, cross-functional leadership, architecture governance, and the ability to work across on-premises and cloud environments.
Deep hands-on expertise in every AI infrastructure technology is not required. The successful candidate should have broad solution-architecture capability, strong infrastructure fundamentals, and sufficient technical depth to guide design decisions, identify dependencies and risks, and collaborate effectively with GPU, network, storage, platform, and software specialists.
Your Role And Responsibilities
AI Data Center Solution Architecture
- Lead requirements discovery and architecture workshops with customer, business, IT, data, infrastructure, security, operations, and vendor stakeholders.
- Translate business, workload, security, performance, availability, and operational requirements into solution blueprints, HLDs, architecture principles, and implementation roadmaps.
- Define the end-to-end architecture across GPU compute, network, storage, management, security, automation, monitoring, platform services, and Day-2 operations.
- Coordinate LLD development with domain architects and engineers, ensuring consistency with the approved HLD, technical standards, and project objectives.
- Define non-functional requirements for scalability, availability, performance, security, recoverability, observability, supportability, and operational readiness.
- Maintain architecture decisions, assumptions, dependencies, technical risks, compatibility requirements, and design-review records.
AI, Data Platform, and Hybrid Infrastructure Integration
- Assess AI and data-platform requirements and map training, inference, validation, development, and tenant workloads to appropriate infrastructure and service models.
- Design integration approaches across on-premises data center, private cloud, and cloud environments, including data flows, control flows, connectivity, security boundaries, and operational dependencies.
- Work with data, AI, software, and infrastructure teams to align platform architecture with workload requirements, data lifecycle, capacity, service levels, and future expansion.
- Support architecture evaluation for multi-tenant resource pools, workload isolation, quotas, reservations, tagging, onboarding, offboarding, and service governance.
- Define architecture validation criteria for MVP, POC, pilot, and production-readiness assessments.
GPU, Network, Storage, and Platform Architecture
- Partner with GPU specialists to define infrastructure principles for NVIDIA GB200, GB300, B200, B300, or equivalent accelerator platforms, including capacity, lifecycle, monitoring, and upgrade considerations.
- Work with network architects to define compute, management, tenant, storage, service, and out-of-band network planes, including segmentation, routing, security, and high-speed connectivity requirements.
- Work with storage specialists to define storage architecture for datasets, models, artifacts, training, inference, checkpoints, backup, restore, and recovery.
- Coordinate technical baselines for operating systems, containers, Kubernetes, drivers, firmware, runtime components, observability, and platform integration.
- Review detailed designs for advanced technologies such as NVLink, InfiniBand, Spectrum-X, RoCE, CUDA, NCCL, and GPU scheduling together with relevant subject-matter experts.
Platform and Automation Integration
- Define integration architecture for cluster lifecycle, tenant management, resource provisioning, network automation, storage services, monitoring, and service orchestration.
- Coordinate architecture requirements for Rafay, Netris, Weka, or comparable platforms, including APIs, workflows, status flows, audit trails, retries, recovery, and rollback.
- Promote API-first integration, reusable templates, automated validation, Infrastructure as Code, and GitOps practices where appropriate.
- Define interface ownership, data exchange, control points, dependency management, and operational handoff between platform components and external systems.
- Ensure architecture supports future automation, service catalog, chargeback/showback, observability, and operational governance requirements.
Customer, Vendor, and Cross-Functional Leadership
- Lead customer workshops, architecture reviews, design walkthroughs, and technical decision meetings in a structured, consultative manner.
- Engage customer executives and technical stakeholders to align business priorities, architecture choices, delivery scope, risks, and roadmap decisions.
- Coordinate internal engineering, product, data, software, and operations teams together with external technology vendors and implementation partners.
- Provide clear architecture status, decision rationale, risk assessments, and recommendations to project and executive stakeholders.
- Support knowledge transfer, solution enablement, and development of reusable architecture patterns and best practices.
Preferred Education
Master's Degree
Required Technical And Professional Expertise
- Degree in Computer Science, Information Technology, Engineering, Telecommunications, Business/Technology Management, or a related field, or equivalent professional experience.
- 10+ years of experience in enterprise technology, telecommunications, infrastructure, digital transformation, solution consulting, product/solution management, or architecture-related roles.
- Demonstrated experience translating business and technical requirements into solution architecture, platform design, implementation roadmaps, or technical-commercial recommendations.
- Experience with enterprise data platforms, AI/big-data initiatives, cloud/on-premises integration, digital transformation, or comparable technology programs.
- Strong understanding of enterprise infrastructure fundamentals, including data centers, networking, IP connectivity, cloud services, security boundaries, and service dependencies.
- Experience leading complex cross-functional or multinational projects and coordinating stakeholders across business, IT, engineering, product, and vendor teams.
- Ability to develop HLDs, architecture blueprints, requirement specifications, decision records, dependency maps, risk assessments, and executive-level technical presentations.
- Strong customer-facing consulting and workshop facilitation skills, including experience engaging senior management and translating between business and technical perspectives.
- Ability to manage changing requirements, competing priorities, delivery risks, and cross-team dependencies in complex project environments.
- Full professional proficiency in Chinese for customer workshops, technical discussions, architecture reviews, and project documentation.
- Working English proficiency for technical documentation, cross-border collaboration, and vendor communication.
- Ability to work locally in Taiwan and provide regular on-site support at customer offices and data center facilities.
Preferred Technical And Professional Experience
- Experience with AI infrastructure, GPU clusters, HPC, private cloud, Kubernetes, or multi-tenant platform architecture.
- Working knowledge of NVIDIA GPU systems and technologies such as NVLink, InfiniBand, Spectrum-X, RoCE, CUDA, NCCL, GPU Operator, or comparable accelerator ecosystems.
- Experience with Rafay, Netris, Weka, or comparable Kubernetes management, network automation, IPAM/SDN, or high-performance storage platforms.
- Experience with API integration, Infrastructure as Code, GitOps, automation frameworks, or platform engineering practices.
- Experience creating detailed LLDs, API/interface specifications, NFR frameworks, compatibility matrices, or production-hardening recommendations.
- Relevant certifications in networking, cloud, Kubernetes, NVIDIA, storage, architecture, or project/program management.