Software Engineer – Cloud Infrastructure at FriendliAI | CA, US | Rezi

Software Engineer – Cloud Infrastructure at FriendliAI

Software Engineer – Cloud Infrastructure

FriendliAI · CA, US

1 months ago

Software Engineer – Cloud Infrastructure

FriendliAI · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Software Engineer – Cloud Infrastructure role.

Rezi rewrites your resume against FriendliAI's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Software Engineer – Cloud Infrastructure posting at FriendliAI — free, in seconds.

About the Role

FriendliAI is seeking a Cloud Infrastructure Engineer to lead the architecture and development of their GPU-accelerated AI inference cloud platform. This role involves designing cluster architecture, extending Kubernetes, and managing the network infrastructure critical for inference traffic.

Responsibilities

  • Own the architecture of the multi-cluster, multi-tenant Kubernetes fleet, including cluster topology, control plane, and etcd lifecycle.
  • Extend Kubernetes using custom controllers, operators, and CRDs.
  • Design GPU scheduling and capacity strategy, encompassing topology-aware placement, node pools, priority, preemption, and tenant quotas.
  • Build autoscaling solutions for inference traffic, including queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.
  • Own the Kubernetes network data plane, including CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.
  • Debug and resolve production network issues, driving permanent fixes.
  • Define SLOs for platform-critical systems and lead post-incident hardening.
  • Deliver infrastructure as code using Terraform, Helm, and GitOps.
  • Collaborate with inference engine, platform, SRE, and security teams to translate serving requirements into platform capabilities.

Requirements

  • 5+ years of experience designing, building, and operating large-scale Kubernetes infrastructure in production.
  • Proven experience operating large-scale, high-traffic network services in production.
  • Deep understanding of Kubernetes internals (API server, scheduler, controller loops, kubelet, etcd).
  • Strong command of Kubernetes and cloud networking (CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, VPC routing).
  • Proficiency with AWS, Terraform, Helm, and Ansible.
  • Programming skills in Go or Python for building infrastructure tooling and automation.
  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.
  • Clear written and verbal communication skills, including the ability to document architectural decisions.

Skills

  • Kubernetes
  • Cloud Networking
  • CNI
  • Kube-proxy/eBPF
  • DNS
  • Load Balancing
  • Service Mesh
  • VPC Routing
  • AWS
  • Terraform
  • Helm
  • Ansible
  • Go
  • Python
  • Distributed Systems Debugging
  • Container Debugging
  • Linux Networking Stack Debugging
  • Cilium
  • eBPF
  • Kubespray
  • NVIDIA GPU Operator
  • RDMA/RoCE
  • InfiniBand
  • EFA
  • SR-IOV
  • NCCL Tuning

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • 5+ years

Education Level

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.

Benefits

  • Flexible working hours
  • Daily lunch and dinner provided
  • Unlimited snacks and beverages
  • Supportive and highly collaborative work environment
  • Health check-up support
  • Top-tier equipment/hardware support
  • Competitive compensation
  • Startup equity
  • Health insurance

About the Company

  • FriendliAI is building the fastest inference cloud for agents, designed to run frontier open-weight models at scale.
  • The platform offers significantly faster output token speed, lower inference costs, and high uptime for demanding agent workloads.
  • It is a fast-moving team focused on generative AI infrastructure.
  • FriendliAI aims to provide a reliable platform for AI inference.