Principal Machine Learning Engineer at Blaze Talent | United States | Rezi

Principal Machine Learning Engineer at Blaze Talent

Principal Machine Learning Engineer

Blaze Talent · United States

Today

Principal Machine Learning Engineer

Blaze Talent · United States

5 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

This is a hands-on, principal-level individual contributor role focused on building and maintaining the systems that turn ML research into a reliable, production-grade product. You will own the pipelines that train models, the evaluation infrastructure, and the serving stack that runs them at scale, making ML industrial-grade.

Responsibilities

  • Build and own training pipelines: data prep, reproducible fine-tuning runs, experiment tracking, and release automation
  • Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning
  • Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management
  • Ship model updates safely with versioning, canarying, rollback, and drift monitoring
  • Build repeatable workflows for adapting models to new domains and customer needs
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets
  • Set the technical bar for ML infrastructure as the team grows

Requirements

  • 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
  • Proven track record owning model serving under real latency, reliability, and cost constraints
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
  • Comfort with high ownership on a small team: scoping your own work, shipping weekly, and making pragmatic build-vs-buy calls
  • Enjoyment of close collaboration with a research counterpart, with clear interfaces and no turf wars

Skills

  • PyTorch
  • distributed training
  • fine-tuning at scale (LoRA, SFT)
  • inference engines such as vLLM or TensorRT-LLM
  • eval harnesses
  • regression gates
  • dataset pipelines
  • precision
  • recall
  • calibration
  • Python
  • containers
  • CI/CD
  • cloud infrastructure
  • observability
  • productionizing small or specialized language models
  • structured-output serving
  • constrained decoding in production

Location

  • New York, NY

Work Type

  • Hybrid

Experience Level

  • Principal

Salary/Compensations

  • $200,000–$250,000 base salary

Benefits

  • Performance bonus
  • meaningful early-stage equity
  • Health, dental, and vision coverage

About the Company

  • We're building AI-native enforcement infrastructure for enterprise communication — technology that catches and fixes compliance issues in real time, before an AI-generated message ever reaches a customer or counterparty, across every channel where AI represents the business.
  • Most existing tools only flag problems after the fact, once the risk is already out the door; we intervene before send.
  • This is a new category, and we're the ones defining it.
  • We're backed by top-tier venture capital and built by a team with backgrounds at major tech and financial firms, led by a founder who has built and scaled AI companies before.