Staff/Senior Distributed Systems Engineer at Helios Intelligence Platforms | NY, US | Rezi

Staff/Senior Distributed Systems Engineer at Helios Intelligence Platforms

Staff/Senior Distributed Systems Engineer

Helios Intelligence Platforms · NY, US

4 days ago

Staff/Senior Distributed Systems Engineer

Helios Intelligence Platforms · NY, US

5 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a distributed systems engineer at Helios, you will expand and operate the distributed execution, storage, scheduling, and reliability primitives that support data ingestion, document processing, search indexing, model inference, agents, and continuously running research workflows. The primary responsibility centers around the orchestration of critical infrastructure to meet mission-critical performance and reliability guarantees. The key objective of this work is to support the real-time access and availability of our entire data corpus as well as supporting the continued construction of the Helios Rapid Ontology System (H.R.O.S.), our long horizon memory data plane.

Responsibilities

  • Own the operation and architecture of Proxi’s distributed execution platform, covering ingestion, document processing, search index performance and long-running research jobs.
  • Manage the deployment of our GovCloud and Air-Gapped resources for sensitive environments.
  • Manage the orchestration of long-running agent research tasks including scheduling, lease management and retention.
  • Queue optimization and cross-cloud information pipeline scalability.
  • Build out dedicated resource-aware autoscaling architecture for fast search and document processing resources.
  • Manage CI/CD and compliance operations across the entire Helios platform.
  • Support global forward embedded customer infrastructure efforts.

Requirements

  • Familiarity with network crawling, compute grids, memory-heavy document processing, OCR, inference, and latency-sensitive jobs.
  • Experience with queue-, log-, workflow-, and actor-based distributed execution systems, including Kafka, Redpanda, Pulsar, SQS, Pub/Sub, RabbitMQ, Temporal, and equivalent technologies.
  • Experience establishing at-least-once delivery with effectively once-only business outcomes through transactional outbox and inbox patterns, sagas, reconciliation, dead-letter handling, replay and historical backfill procedures.
  • Experience with CPU-, memory-, disk-, network-, and GPU-aware scheduling, including workload classification, priority allocation, starvation prevention, tenant and dependency concurrency limits, placement constraints, resource quotas, and noisy-neighbor isolation.
  • Experience with PostgreSQL, object storage, Redis or equivalent caches, search indices (Elasticsearch or Typesense), vector stores, change-data-capture, and graph-storage systems.
  • Experience with distributed coordination and concurrency controls, including leader election, locks, leases, fencing tokens, optimistic concurrency, conflict resolution, and partition recovery.
  • Experience with provisioning and operation of AWS, GCP, or Azure infrastructure through terraform or similar IaC, vulnerability scanning, CI/CD and deployment methods.
  • Experience with definition and administration of SLO/Is, error budgets, release controls, metrics, OpenTelemetry, Datadog, fault injection and load/failure testing.
  • Experience with evaluation of system performance and cost via e2e latency decomposition, resource and dependency profiling.
  • Experience with enforcement of platform security and tenant isolation policies.
  • Ability to own outcomes, not just assigned tasks.
  • Ability to move quickly without lowering the standard.
  • Ability to stay close to the mission and the user.
  • Ability to work across boundaries.
  • Ability to communicate directly.
  • Ability to build for the real world.
  • Flexibility during critical deployment periods, customer incidents, product launches, and other company-critical work.

Skills

  • Distributed execution
  • Storage
  • Scheduling
  • Reliability primitives
  • Data ingestion
  • Document processing
  • Search indexing
  • Model inference
  • Agents
  • Research workflows
  • Orchestration
  • Kafka
  • Redpanda
  • Pulsar
  • SQS
  • Pub/Sub
  • RabbitMQ
  • Temporal
  • Transactional outbox
  • Inbox patterns
  • Sagas
  • Reconciliation
  • Dead-letter handling
  • Replay
  • Historical backfill
  • CPU-aware scheduling
  • Memory-aware scheduling
  • Disk-aware scheduling
  • Network-aware scheduling
  • GPU-aware scheduling
  • Workload classification
  • Priority allocation
  • Starvation prevention
  • Tenant concurrency limits
  • Dependency concurrency limits
  • Placement constraints
  • Resource quotas
  • Noisy-neighbor isolation
  • PostgreSQL
  • Object storage
  • Redis
  • Elasticsearch
  • Typesense
  • Vector stores
  • Change-data-capture
  • Graph-storage systems
  • Distributed coordination
  • Concurrency controls
  • Leader election
  • Locks
  • Leases
  • Fencing tokens
  • Optimistic concurrency
  • Conflict resolution
  • Partition recovery
  • AWS
  • GCP
  • Azure
  • Terraform
  • Infrastructure as Code (IaC)
  • Vulnerability scanning
  • CI/CD
  • Deployment methods
  • SLO/Is
  • Error budgets
  • Release controls
  • Metrics
  • OpenTelemetry
  • Datadog
  • Fault injection
  • Load testing
  • Failure testing
  • E2E latency decomposition
  • Resource profiling
  • Dependency profiling
  • Platform security
  • Tenant isolation policies
  • GovCloud
  • Air-Gapped resources
  • Autoscaling
  • Compliance operations
  • Natural language understanding
  • Personalized relevance

Location

  • New York City

Work Type

  • Full-time
  • On-site

Experience Level

  • Staff
  • Senior

About the Company

  • Helios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer.
  • Government shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer.
  • Our core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments.
  • We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry.
  • Helios is a fast-moving startup with ambitious goals.
  • This is not a conventional 9-to-5 role.
  • This role offers unusual ownership, direct access to consequential institutions, and the opportunity to build systems that affect how major decisions are made.

Equal Opportunity

  • Helios is an equal opportunity employer.
  • We evaluate candidates based on their abilities, experience, and potential to contribute to our mission.
  • We do not discriminate on the basis of race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other status protected by applicable law.