Platform / Site Reliability Engineer at Sunset | NY, US | Rezi

Platform / Site Reliability Engineer at Sunset

Platform / Site Reliability Engineer

Sunset · NY, US

1 months ago

Platform / Site Reliability Engineer

Sunset · NY, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Platform / Site Reliability Engineer role.

Rezi rewrites your resume against Sunset's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Platform / Site Reliability Engineer posting at Sunset — free, in seconds.

About the Role

Sunset operates customer-facing SaaS products, connector and ingestion services, asynchronous workers, high-volume data pipelines, model-backed systems, review tools, and customer-delivery paths. You will build and operate the shared platform that lets our product, data, and AI teams ship reliable, secure, observable, and cost-aware systems without manual infrastructure work or operational risk growing linearly. This is not a deployment-operator or internal-IT role. You will give product, data, and ML teams the runtime, delivery, visibility, recovery, and operating patterns to own their systems well.

Responsibilities

  • Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational tooling
  • Create supported paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloads
  • Improve deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediation
  • Work with engineering teams to define useful service and pipeline objectives, ownership, escalation, and recovery paths
  • Build self-service for common infrastructure, environment, access, deploy, debugging, and recovery work
  • Make cloud and vendor cost understandable by service and workload and improve efficiency within explicit reliability and security bounds
  • Partner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidence
  • Support employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authority
  • Use AI tools deeply for platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusions

Requirements

  • You have personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including an asynchronous, batch-data, or model-backed workload
  • You are a strong software engineer who is comfortable changing application, platform, and infrastructure code and operating the result in production
  • You can reason from user impact through dependencies, state, telemetry, incident response, recovery, and durable remediation
  • You have built paved roads other engineers adopted because they made real work easier, not because a platform team required them
  • You understand both long-running services and high-volume or scheduled workloads and know where their reliability models should differ
  • You can make pragmatic tradeoffs among delivery speed, least privilege, isolation, recovery, developer experience, and unit cost
  • You are effective in an early-stage environment where the first step is often to establish ownership and a trustworthy baseline
  • You can lead calmly through ambiguous incidents, communicate clearly, and leave the system and operating model stronger afterward
  • You use modern AI engineering tools fluently and verify generated infrastructure, queries, code, and operational conclusions before they affect production

Skills

  • AWS
  • Terraform
  • container runtimes
  • workflow orchestration
  • observability systems
  • high-volume data processing
  • model serving
  • evaluation jobs
  • GPU workloads
  • machine-learning platforms
  • developer environments
  • preview systems
  • CI/CD
  • progressive delivery
  • internal developer platforms
  • replayable pipelines
  • backup and restore
  • disaster recovery
  • capacity planning
  • cloud-cost allocation
  • technical controls
  • automated evidence
  • SOC 2

Experience Level

  • early platform or SRE hire at a fast-growing company

About the Company

  • Sunset was founded to help founders, initially supporting startups through shutting down, and has expanded into unlocking a new revenue stream for all types of businesses.
  • The company's insight in 2025 was that the data generated daily through collaboration, communication, and building is valuable training data for AI models.
  • Sunset partners directly with frontier AI labs, providing a primary source of real, proprietary data.
  • The company has scaled from $0 to a multi-eight-figure run rate in months.
  • Sunset has raised funding from investors including Floodgate, Afore, Ludlow, and Hustle Fund.