About the Role
This is a hands-on, principal-level individual contributor role focused on building and maintaining the systems that turn ML research into a reliable, production-grade product. You will own the pipelines that train models, the evaluation infrastructure, and the serving stack that runs them at scale, making ML industrial-grade.
Responsibilities
- Build and own training pipelines: data prep, reproducible fine-tuning runs, experiment tracking, and release automation
- Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning
- Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management
- Ship model updates safely with versioning, canarying, rollback, and drift monitoring
- Build repeatable workflows for adapting models to new domains and customer needs
- Convert expert labels and reviewer feedback into clean training and evaluation datasets
- Set the technical bar for ML infrastructure as the team grows
Requirements
- 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
- Proven track record owning model serving under real latency, reliability, and cost constraints
- Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
- Comfort with high ownership on a small team: scoping your own work, shipping weekly, and making pragmatic build-vs-buy calls
- Enjoyment of close collaboration with a research counterpart, with clear interfaces and no turf wars
Skills
- PyTorch
- distributed training
- fine-tuning at scale (LoRA, SFT)
- inference engines such as vLLM or TensorRT-LLM
- eval harnesses
- regression gates
- dataset pipelines
- precision
- recall
- calibration
- Python
- containers
- CI/CD
- cloud infrastructure
- observability
- productionizing small or specialized language models
- structured-output serving
- constrained decoding in production
Location
- New York, NY
Work Type
- Hybrid
Experience Level
- Principal
Salary/Compensations
- $200,000–$250,000 base salary
Benefits
- Performance bonus
- meaningful early-stage equity
- Health, dental, and vision coverage
About the Company
- We're building AI-native enforcement infrastructure for enterprise communication — technology that catches and fixes compliance issues in real time, before an AI-generated message ever reaches a customer or counterparty, across every channel where AI represents the business.
- Most existing tools only flag problems after the fact, once the risk is already out the door; we intervene before send.
- This is a new category, and we're the ones defining it.
- We're backed by top-tier venture capital and built by a team with backgrounds at major tech and financial firms, led by a founder who has built and scaled AI companies before.
