Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Staff ML Platform Engineer (MLOps) role.
Rezi rewrites your resume against FutureFit AI's job description. Free.

Tailor your resume to this Staff ML Platform Engineer (MLOps) role.
Rezi rewrites your resume against FutureFit AI's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Staff ML Platform Engineer (MLOps) posting at FutureFit AI — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Staff ML Platform Engineer (MLOps) posting at FutureFit AI — free, in seconds.
About the Role
We're seeking a Staff ML Platform Engineer (MLOps) to build and operate the platform our ML and LLM-powered products run on. You will own the end-to-end lifecycle of models, from building and deployment to evaluation and serving, including compute and environment management, LLM call routing and optimization, and system monitoring.
Responsibilities
- Evaluate current pipelines, data architecture, and ML workflows to create a prioritized plan for improvement.
- Own compute provisioning, environment management, and ensure reproducible training and serving environments.
- Tune latency, throughput, and spend without destabilizing production.
- Build the infrastructure for LLM features, including smart routing for cost-efficiency and prompt/response evaluation.
- Implement A/B testing for models before full rollout, including shadow deploys, canaries, and holdouts with agreed-upon success criteria.
- Develop model and data monitoring, alerting, regression detection, lineage, and traceability to explain and reproduce predictions.
- Operate running systems, debug production ML incidents, and fix systems including those built by others.
- Own feature computation, storage, and serving consistency between training and inference, including necessary data engineering.
- Establish deployment and monitoring standards for the team to adopt, fostering shared ownership.
Requirements
- Staff-level, hands-on experience in MLOps, ML platform, or ML infrastructure.
- Experience standing up MLOps practice end to end: CI/CD for models, experiment tracking, model registries, deployment workflows, and monitoring.
- Production experience with LLM-based systems: serving, prompt and response evaluation, routing across models and providers, and managing cost and latency tradeoffs.
- Experience operating models in both batch and real-time serving contexts.
- Hands-on with compute provisioning and environment management: containers, reproducible training and serving environments, and keeping frameworks and packages current.
- Experience running controlled model experiments in production: A/B tests, shadow or canary deploys, holdouts, and setting success criteria.
- Depth in observability and traceability for production ML: drift and regression detection, alerting, lineage, and ability to trace predictions.
- Hands-on operational experience: carrying the pager or equivalent, debugging production ML incidents under pressure, and fixing systems not originally built.
- Comfort doing data engineering for the platform: pipelines, feature computation and storage, and keeping training and serving features consistent.
- Track record of diagnosing problems and materially improving complex, fast-grown systems.
- Strong systems design ability: translate product needs into durable architecture and build it yourself.
Skills
- MLOps
- ML platform
- ML infrastructure
- Data Platform Engineer
- ML Infrastructure Engineer
- Data Scientist
- CI/CD for models
- Experiment tracking
- Model registries
- Deployment workflows
- Monitoring
- LLM-based systems
- Prompt and response evaluation
- Routing across models and providers
- Cost and latency tradeoffs
- Batch serving
- Real-time serving
- Compute provisioning
- Environment management
- Containers
- Reproducible training and serving environments
- Frameworks and packages management
- A/B testing
- Shadow deploys
- Canary deploys
- Holdouts
- Observability
- Traceability
- Drift detection
- Regression detection
- Alerting
- Lineage
- Data engineering
- Pipelines
- Feature computation
- Feature storage
- Systems design
Location
- Remote (CA/US)
- Toronto (optional office)
Work Type
- Remote-first
- Full-time
Experience Level
- Staff-level
Education Level
- No specific degree required; focus on grit, hunger, drive, continuous learning, and tackling challenges.
Salary/Compensations
- USD $172,000-$215,000 (United States)
- CAD $172,000-$220,000 (Canada)
Benefits
- Company off-site
About the Company
- FutureFit AI's mission is to help more people get to better jobs faster and cheaper, focusing on those facing barriers to opportunity and addressing economic inequality.
- Our AI-powered platform brings efficiency and insight to workforce development.
- We are a remote-first company with hubs in NYC and Toronto.
- Our industry is SaaS/AI technology.
- We are bootstrapped and have raised funding led by JP Morgan.
- Our core principles are: Be Curious, Drive to Outcomes, Raise the Bar, Speed Matters, Own It, We Over Me.
Equal Opportunity
- We are proud to be an equal opportunity workplace.
- We celebrate diversity and are committed to creating an inclusive environment for all employees.
- We do not discriminate on the basis of race, religion, color, gender identity, sexual orientation, age, disability, veteran status, or other applicable legally protected characteristics.
- We encourage people of different backgrounds, experiences, abilities, and perspectives to apply.
- We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive other benefits and privileges of employment.