About the Role
As a Senior Data Engineer on the AI Systems team, you will own the Spark-based data pipelines and data infrastructure that power our core recommendations engine and ML platform. You will build, scale, and optimize the data layer feeding our production ML models, working collaboratively with ML engineers and scientists on large-scale data systems that impact millions of customers.
Responsibilities
- Build, maintain, and optimize production data pipelines for AI-driven personalization.
- Own and scale Spark-based batch pipelines, including cluster configuration, tuning, and performance optimization on GCP Dataproc.
- Build and maintain the ML Data Lake, ensuring data quality, accessibility, and efficient storage.
- Support the data needs of ML Engineers and Scientists for model development, training, and evaluation.
- Identify and resolve performance bottlenecks and scaling limitations in data pipelines and infrastructure.
- Collaborate with distributed systems engineers on architectural evolution, ensuring data layer continuity.
- Continuously improve data infrastructure for scalability and reliability.
- Release features and data products that deliver measurable business value.
Requirements
- 5+ years of data engineering experience.
- Deep expertise with Apache Spark, including the PySpark DataFrame API and experience solving challenging scaling problems.
- Experience with large-scale data processing, cluster configuration, optimization, and tuning (GCP Dataproc).
- Strong software development skills in Python (unit testing, git, code review, CI/CD).
- Experience with data storage formats (Parquet, Delta Lake).
- Experience with event streaming data (Kafka).
- Experience with cloud computing platforms (Google Cloud Platform).
- Experience with advanced query optimization.
- Familiar with Software Development Lifecycle practices, such as continuous integration/continuous delivery and automated deployment (Docker, Kubernetes, GitHub Actions).
- Ability to collaborate with technical partners to determine requirements and make design decisions.
- Enjoys working in a fast-paced, goal-driven environment.
Skills
- Apache Spark
- PySpark DataFrame API
- GCP Dataproc
- Python
- Unit testing
- Git
- Code review
- CI/CD
- Parquet
- Delta Lake
- Kafka
- Google Cloud Platform
- Docker
- Kubernetes
- GitHub Actions
Location
- New York City
Work Type
- Full-time
Experience Level
- Senior
Salary/Compensations
- 144K CAD - 188K CAD
Benefits
- Additional bonus
- Full range of medical, financial, and/or other benefits
About the Company
- Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning.
- The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility.
- Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.
Equal Opportunity
- We are committed to building a diverse and inclusive culture where all Inkers can thrive.
- We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics.
- We are proud to be an equal opportunity employer.
