Senior Software Engineer, Machine Learning Platform at Chime Financial, Inc | CA, US | Rezi

Senior Software Engineer, Machine Learning Platform at Chime Financial, Inc

Senior Software Engineer, Machine Learning Platform

Chime Financial, Inc · CA, US

1 weeks ago

Senior Software Engineer, Machine Learning Platform

Chime Financial, Inc · CA, US

12 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Software Engineer, Machine Learning Platform role.

Rezi rewrites your resume against Chime Financial, Inc's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Software Engineer, Machine Learning Platform posting at Chime Financial, Inc — free, in seconds.

About the Role

Chime’s Machine Learning Platform (MLP) team builds and operates the infrastructure, tooling, and developer experience that powers machine learning across the company. This role focuses on creating secure, reliable, and reusable platform capabilities that help teams choose the right approach, from conventional predictive models to LLM-powered and multi-step agentic systems, while maintaining strong standards for evaluation, observability, governance, privacy, and cost efficiency. You will design and build scalable systems spanning traditional machine learning and emerging AI workloads, including model training, feature computation, real-time inference, foundation-model access, evaluation, and agentic orchestration.

Responsibilities

  • Design, build, and operate scalable ML and AI infrastructure on AWS.
  • Design and operate shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
  • Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
  • Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
  • Help teams make principled architecture decisions across traditional ML, LLM-powered applications, and agentic workflows, and contribute to the platform’s technical roadmap.
  • Build distributed training, batch inference, and large-scale processing systems using frameworks such as Ray or Spark.
  • Build and maintain infrastructure as code using Terraform.
  • Support and evolve the feature store and feature pipelines.
  • Develop data ingestion and streaming systems using technologies such as Kinesis, Kafka, Flink, or Spark.
  • Improve CI/CD workflows for ML models, AI applications, and platform components.
  • Partner closely with Data Science and ML Engineering teams to improve developer experience.
  • Participate in on-call rotations to support production systems.

Requirements

  • Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
  • Experience designing distributed systems and large-scale data or compute platforms on AWS using frameworks such as Spark or Ray.
  • 5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
  • Working knowledge of LLM application patterns such as retrieval-augmented generation, structured outputs, tool calling, agent orchestration, and evaluation of non-deterministic systems.
  • Experience designing production systems that integrate ML or foundation models through reliable APIs, workflows, and data contracts.
  • Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Strong programming skills in Python, Go, Scala, Java, or similar languages.
  • Solid understanding of software engineering fundamentals, including testing, version control, code review, and observability.
  • Experience shipping LLM-powered or agentic systems to production.
  • Experience with one or more of the following: model gateways, prompt lifecycle management, retrieval or vector search, tool execution, and agent orchestration frameworks.
  • Experience building evaluation, tracing, and observability capabilities for non-deterministic AI systems.
  • Familiarity with managed or self-hosted foundation model infrastructure, such as Amazon Bedrock, SageMaker, or equivalent platforms.
  • Experience operating GPU-based workloads and optimizing training or inference performance and cost; CUDA experience is a plus.

Skills

  • AWS
  • Spark
  • Ray
  • Terraform
  • Kinesis
  • Kafka
  • Flink
  • Docker
  • Kubernetes
  • Python
  • Go
  • Scala
  • Java
  • LLM
  • Agentic systems
  • Retrieval-augmented generation
  • Structured outputs
  • Tool calling
  • Agent orchestration
  • Evaluation of non-deterministic systems
  • CI/CD
  • DevOps
  • Infrastructure as code
  • Observability
  • Model gateways
  • Prompt lifecycle management
  • Retrieval
  • Vector search
  • Tool execution
  • Amazon Bedrock
  • SageMaker
  • GPU
  • CUDA

Location

  • Onsite
  • Remote

Work Type

  • Full-time
  • Hybrid

Experience Level

  • Senior
  • 5+ years

Salary/Compensations

  • $187,000.00 - $259,000.00

Benefits

  • Bonus
  • Competitive equity package
  • Backup child, elder, and pet care
  • Subsidized commuter benefits
  • Comprehensive health, financial, and wellbeing benefits
  • Generous vacation policy
  • Company-wide paid days off
  • 1% of your time off to support local community organizations
  • Annual wellness stipend
  • Up to 22 weeks of paid parental leave for birthing parents
  • 12 weeks of paid parental leave for non-birthing parents
  • Access to family planning reimbursement

About the Company

  • At Chime, we believe that everyone can achieve financial progress. We created Chime—a financial technology company, not a bank*—on the premise that core banking services should be helpful, easy, and free. Through our user-friendly tools and intuitive platforms, we empower our members to take control of their finances and work towards their goals. We're a team of problem solvers, dreamers, and builders with one shared obsession: our members. We believe in being bold, dreaming big, and taking risks, while also working together, embracing our diverse perspectives, and giving each other honest feedback. Our culture remains deeply entrepreneurial, encouraging every Chimer to see themselves as stewards of our mission to help everyday Americans unlock their financial progress. We know that to achieve our mission, we must earn and keep people's trust—so we hold ourselves to the highest standards of integrity in everything we do.

Equal Opportunity

  • Chime is proud to be an Equal Opportunity Employer. We consider qualified applicants without regard to race, color, ancestry, religion, sex, national origin, sexual orientation, gender identity, age, marital or family status, disability, genetic information, veteran status, or any other legally protected basis under provincial, federal, state, and local laws, regulations, or ordinances. We will also consider qualified applicants with criminal histories in a manner consistent with the requirements of state and local laws, including the San Francisco Fair Chance Ordinance, Cook County Ordinance, NYC Fair Chance Act, and the LA City Fair Chance Ordinance, and consistent with Canadian provincial and federal laws. If you have a disability or special need that requires accommodation during any stage of the application process, please contact: accommodations@chime.com.