About the Role
We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. This role advises on how AI services and agents are built and deployed, how models execute, how requests are routed, how AI assets are catalogued and governed, how consumption is measured, and how the platform is observed across various environments. This role is ideal for an engineer comfortable building automated delivery pipelines, operating high-performance inference, and building deep observability and cost visibility, understanding that AI workloads must be delivered repeatably and requests economically bounded, attributable, and traceable.
Responsibilities
- Supports DevOps and delivery for AI workloads, building and operating CI/CD/CV pipelines for AI services, agents, and runtime components.
- Owns governance and discovery for AI assets, including service catalog/registry, experiment tracking, model metadata, upstream registries/mirrors, CVE/SBOM scanning, lineage contracts, and license management.
- Owns resource and cost management, including quotas, rate limits, cost attribution, and utilization.
- Owns the full observability stack, including metrics, logs, traces, dashboards, LLM debugging and evaluation, and SLA/alert notifications.
- Owns the OpenTelemetry collection layer, including multi-tenant receiver, exporters, queues, DCGM exporter for GPU telemetry, processor batching, and dynamic filtering.
- Automates GitOps-based delivery and continuous verification, embedding quality, integrity, and cost gates into pipelines.
- Closes the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.
- Ensures cost and telemetry are identity-stamped and per-tenant for end-to-end attribution and tying FinOps and observability to workloads.
Requirements
- 8+ years in DevOps, MLOps, platform, or observability engineering, with hands-on production ownership of AI or high-throughput services.
- Strong hands-on DevOps experience, including CI/CD/CV pipelines and GitOps tooling for automated build, test, release, and rollback.
- Hands-on expertise operating inference/model-serving frameworks on GPU infrastructure.
- Strong experience with observability stacks and OpenTelemetry.
- Experience with API gateways and request routing, including streaming responses.
- Experience with cost management / FinOps tooling and quota/rate-limit enforcement.
- Familiarity with model/artifact registries and supply-chain scanning.
- Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.
- Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams.
Skills
- Strong DevOps expertise: CI/CD/CV pipeline design, GitOps, continuous verification, and progressive/automated release and rollback for production workloads.
- Deep expertise operating model-serving and inference systems on GPUs at production scale.
- Deep observability skills: metrics, logs, traces, and OpenTelemetry.
- FinOps mindset: able to attribute, bound, and optimize AI consumption cost per tenant and workload.
- Familiarity with model/artifact governance, registries, CVE scanning, and license/lineage tracking.
- Comfortable operating across cloud, on-prem, edge, and air-gapped environments with consistent runtime and telemetry semantics.
- Strong communicator able to explain runtime, cost, and observability tradeoffs to engineers, architects, and leadership.
Location
- Anywhere in Country
Work Type
- Hybrid
Experience Level
- 8+ years
Education Level
- Bachelor's or Master's degree in Computer Science or related technical field
Salary/Compensations
- US: $106,900 to $176,500
- New York City Metro Area, Washington State and California (excluding Sacramento): $128,400 to $200,600
Benefits
- Comprehensive compensation and benefits package
- Medical and dental coverage
- Pension and 401(k) plans
- Wide range of paid time off options
- Flexible vacation policy
- Time off for designated EY Paid Holidays, Winter/Summer breaks, Personal/Family Care, and other leaves of absence
About the Company
- EY is building a better working world by creating new value for clients, people, society and the planet, while building trust in capital markets.
- Enabled by data, AI and advanced technology, EY teams help clients shape the future with confidence and develop answers for the most pressing issues of today and tomorrow.
- EY teams work across a full spectrum of services in assurance, consulting, tax, strategy and transactions.
- Fueled by sector insights, a globally connected, multi-disciplinary network and diverse ecosystem partners, EY teams can provide services in more than 150 countries and territories.
Equal Opportunity
- EY provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, genetic information, national origin, protected veteran status, disability status, or any other legally protected basis, including arrest and conviction records, in accordance with applicable law.
- EY is committed to providing reasonable accommodation to qualified individuals with disabilities including veterans with disabilities.
