We are seeking a Senior Backend Engineer with passion to bring our products to success.
The ideal candidate will support the AI agent platform backend team by reviewing system architecture design, and spearheading problem-solving initiatives.
<Responsibilities>
- Design and maintain low-latency, high-availability streaming services and event protocols that support real-time integration between client applications and AI agentic services at scale.
- Architect and optimize state management and retrieval across relational and NoSQL document stores to support long-term AI memory, including vector-based semantic retrieval.
- Reduce LLM token cost through prompt restructuring, caching of repeated context, and consolidating the number of model calls per conversation turn.
- Diagnose and resolve Kubernetes resource bottlenecks under high concurrency, covering resource allocation, health checks, and node pool planning.
- Deploy distilled models on GPU node pools, and manage model routing, load balancing, rate limiting, and failover through an LLM gateway.
<Required Skills>
- Bachelor's degree in Computer Science or equivalent practical experience.
- 2+ years of experience in backend infrastructure, API design, and distributed systems development.
- Experience with Python, including asynchronous programming, streaming response handling, connection pooling, and debugging event loop behavior in production.
- Experience operating relational databases, distributed caching layers, and containerized services in production.
- Experience operating Kubernetes in production.
- Experience building, deploying, and maintaining high-concurrency services on GCP or AWS.
- Experience running managed Kubernetes, managed relational databases, message queues, container registries, and centralized logging in production, with a track record of reducing cloud infrastructure cost.
- Experience operating LLM applications in production, including latency and cost governance, caching strategies, streaming responses, and tool/function calling.
- Experience implementing retrieval-augmented generation systems or vector databases at scale, with an understanding of vector similarity search trade-offs.
- Experience designing or implementing agentic workflows, multi-agent systems, or the Model Context Protocol.
- Experience serving self-hosted models on GPU infrastructure, including scheduling, and deploying quantized or distilled models with high-throughput inference engines.