Job Description
Job Summary
We are seeking a
Principal Core Infrastructure Engineer who is passionate about solving complex engineering problems in distributed systems, virtualized infrastructure, and highly available storage services.
In this hands-on technical role, you will lead the design and development of major components and features within the OCI Block Storage Service. You will architect scalable and elastic distributed systems, optimize critical code and data paths, and design systems that maintain correctness, durability, availability, and performance at hyperscale.
You will also play an important role in operational excellence by improving observability, diagnosing complex production issues, strengthening security and compliance, driving automation, and mentoring other engineers.
Success in this role requires strong programming fundamentals, deep distributed systems expertise, sound architectural judgment, and the ability to work effectively across both detailed implementation and broad system-level design.
Responsibilities
Key Responsibilities
- Design and architect scalable distributed systems: Lead the design and implementation of major Block Storage components that support horizontal and vertical scaling, elasticity, distributed state management, and large-scale data processing.
- Drive performance and efficiency at scale: Optimize software, data paths, and system interactions for high-throughput, hyperscale workloads; establish scalability requirements and develop appropriate performance and load-testing strategies.
- Engineer highly reliable and durable services: Design fault-tolerant systems using redundancy, replication, synchronization, and automated failover while enabling in-service upgrades and minimizing customer-visible disruption.
- Design for distributed-systems failure modes: Develop systems that behave predictably during network partitions, infrastructure failures, and traffic spikes by applying appropriate consistency and availability models, load shedding, throttling, and rate limiting.
- Ensure correctness, availability, and observability: Define functional and correctness requirements, establish KPIs and telemetry, build dashboards and alerts, and create sophisticated validation strategies such as fault injection and brownout testing.
- Own operational excellence: Proactively diagnose, debug, and resolve complex production issues; participate in operational support rotations; conduct or guide root-cause investigations; and ensure systems meet operational-readiness standards.
- Strengthen security and compliance: Implement security controls for multi-tenant cloud environments, including encryption and access controls; remediate identified vulnerabilities and maintain infrastructure and supporting documentation in accordance with applicable requirements.
- Advance automation and engineering practices: Develop Infrastructure as Code (IaC), tooling, and automated workflows for provisioning, patching, updating, deployment, and rollback while mentoring engineers and contributing to strong engineering practices across the team.
Qualifications & Skills Mandatory
- 10+ years of software engineering experience, including experience designing and delivering large-scale, highly available distributed systems or backend services.
- Strong hands-on programming experience developing production-quality software in Java, C++, or comparable systems/backend programming languages.
- Experience with scripting or automation languages such as Python.
- Strong knowledge of data structures, algorithms, concurrency, multithreading, and distributed systems fundamentals.
- Demonstrated experience designing systems for scalability, high availability, fault tolerance, durability, and performance.
- Working knowledge of networking fundamentals and protocols, including TCP/IP, HTTP, and common network architectures.
- Strong production troubleshooting, debugging, performance analysis, and performance-tuning capabilities.
- Ability to independently own complex technical components while collaborating effectively in an agile engineering environment.
Good-to-Have
- Experience building or operating multi-tenant cloud infrastructure.
- Experience with virtualized storage infrastructure, block storage, storage systems, or data-management platforms.
- Experience designing replication, synchronization, consistency, and failover mechanisms in distributed systems.
- Experience with cloud service observability, including telemetry, metrics, dashboards, alerting, and SLO-based operations.
- Experience designing and executing fault-injection, brownout, resilience, performance, or large-scale load tests.
- Experience with Infrastructure as Code (IaC) and automated cloud infrastructure provisioning, deployment, patching, and rollback.
- Knowledge of cloud security practices, including encryption, access controls, vulnerability remediation, and compliance requirements.
- Experience participating in production support rotations, incident response, and root-cause analysis for large-scale services.
- Demonstrated ability to mentor engineers, influence technical direction, and contribute to architectural decisions across teams.
Self-Assessment Questions
Before applying, consider the following:
- Have I spent significant time designing and delivering large-scale, highly available distributed systems or backend services in production environments
- Can I independently design a distributed-system component while reasoning about scalability, concurrency, fault tolerance, consistency, durability, availability, and performance trade-offs
- Do I have strong hands-on coding experience in Java, C++, or a comparable language, along with solid knowledge of algorithms, data structures, and multithreaded programming
- Can I diagnose complex production issues across software, networking, infrastructure, and distributed-system interactions and drive them through root-cause analysis and resolution
- Am I comfortable owning the full lifecycle of a critical service component—from architecture and implementation through testing, observability, security, deployment, and production operations
Qualifications
Career Level - IC4
About Us
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing [Confidential Information] or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.