Senior Platform Reliability Engineer (Fabric and Interconnect) at Firmus Technologies | AU | Rezi

Senior Platform Reliability Engineer (Fabric and Interconnect) at Firmus Technologies

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies · AU

Yesterday

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies · AU

a day ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Platform Reliability Engineer (Fabric and Interconnect) role.

Rezi rewrites your resume against Firmus Technologies's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Platform Reliability Engineer (Fabric and Interconnect) posting at Firmus Technologies — free, in seconds.

About the Role

Firmus Technologies is seeking a Senior Platform Reliability Engineer, Fabric and Interconnect to own the reliability of the estate's GPU interconnect and network fabrics. This hands-on role involves deep technical expertise and a strong focus on automation to ensure the self-healing capability of the infrastructure.

Responsibilities

  • Responsible for the reliable operation, automation, and continuous improvement of the estate's GPU interconnect and network fabrics.
  • Build and maintain guarded automation and remediation tooling for fabric faults.
  • Diagnose and tune performance across the interconnect stack.
  • Operate DPU-based host networking across the fleet.
  • Execute firmware upgrade waves, fabric expansions, and capacity changes.
  • Provide deep technical expertise for fabric and interconnect faults.
  • Lead vendor escalations at an engineering level.
  • Lead technical recovery during major fabric incidents.
  • Share the follow-the-sun on-call roster.
  • Mentor engineers carrying frontline diagnosis.
  • Document operational procedures, runbooks, and performance results.

Requirements

  • 8+ years of experience in high-performance networking and systems engineering.
  • Substantial ownership of production network or interconnect infrastructure in a 24/7 environment.
  • Deep operational experience with high-performance GPU interconnect fabrics.
  • Extensive experience with high-performance networking fabrics.
  • Experience with DPU or SmartNIC-based host networking.
  • Experience planning and executing firmware upgrade waves and capacity expansions on production fabrics.
  • Strong skills in infrastructure automation and infrastructure-as-code practices.
  • Practical experience with scripting or programming for operational automation and tooling (e.g., Python, Go, Bash).
  • Proven ability to act as a senior escalation point in production.
  • Clear technical judgment and communication skills.

Skills

  • High-performance networking
  • Systems engineering
  • GPU interconnect fabrics (NVLink, NVSwitch)
  • High-performance networking fabrics (InfiniBand, RoCE-based Ethernet, NVIDIA Spectrum-X)
  • DPU or SmartNIC-based host networking
  • Infrastructure automation
  • Infrastructure-as-code
  • Scripting/programming (Python, Go, Bash)
  • Major incident response
  • Vendor escalation
  • Runbook production

Location

  • Australia
  • Singapore

Work Type

  • Permanent
  • Full-time
  • On-call

Experience Level

  • Senior

About the Company

  • Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
  • Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
  • We design, build, and operate a new class of digital infrastructure – the AI Factory.
  • Our model-to-grid technology approach pushes the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction.
  • Firmus AI Cloud is a large-scale GPU cloud platform purpose-built to deliver energy-efficient AI compute at scale.
  • AI FactoryOS is Firmus' proprietary operating system for the AI Factory, governing GPU telemetry, cooling, power, and grid interaction.
  • The AI FactoryOS Operations function runs the platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud, and the platforms built on them.

Equal Opportunity

  • At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.