About the Role
Engage with complex business challenges, harnessing modern tools and technologies to securely store, process, transform, and enrich terabyte to petabyte scale healthcare data. Your work will underpin data-driven business decisions and contribute to our mission of delivering industry-best data products / software with a customer-first mindset and team-oriented approach.
Responsibilities
- Build and optimize analytical data models in BigQuery, implementing partitioning, clustering, and materialized views for performance and cost efficiency.
- Ensure compliance with data governance, access controls, and IAM best practices.
- Develop integrations with external systems (APIs, flat files etc.) using GCP-native or hybrid approaches.
- Utilize tools like Dataflow or custom Python/Java services on Cloud Functions or Cloud Run to handle transformations and ingestion logic.
- Build automated CI/CD pipelines using Cloud Build, GitHub Actions, or Jenkins for deploying data pipeline code and workflows.
- Set up observability using Cloud Monitoring, Cloud Logging, and Error Reporting to ensure pipeline reliability.
- Design, develop, and maintain optimal data pipelines to assemble large and intricate datasets, catering to the business requirements of various CVS lines of business.
- Craft tools to provide actionable insights and integrate them with consumer touchpoints.
- Collaborate with cross-functional teams in a dynamic and agile environment.
- Solve problems associated with large scale complex, structured and unstructured data.
Requirements
- 1+ years of experience with SQL, NoSQL
- 1+ years of experience with Python (or a comparable scripting language)
- 1+ years of experience with Data warehouses (such as data modeling and technical architectures) and infrastructure components
- 1+ years of experience with ETL/ELT, and building high-volume data pipelines
- 1+ years of experience with reporting/analytic tools
- 1+ years of experience with query optimization, data structures, transformation, metadata, dependency, and workload management
- 1+ years of experience with Big data and cloud architecture
- 1+ years of hands-on experience building modern data pipelines within a major cloud platform (GCP)
- 1+ years of experience with deployment/scaling of apps on containerized environment (i.e. Kubernetes, AKS)
- 1+ years of experience with real-time and streaming technology (i.e. Kafka, Spark Streaming)
- 1+ year(s) of soliciting complex requirements and managing relationships with key stakeholders
- Experience with complex systems and solving challenging analytical problems
- Strong collaboration and communication skills within and across teams
- Knowledge of data visualization and reporting
- Experience in designing and building data engineering solutions in cloud environments (preferably GCP)
- Experience with Git, CI/CD pipeline, and other DevOps principles/best practices
- Experience with bash shell scripts, UNIX utilities & UNIX Commands
- Understanding of software development methodologies including waterfall and agile
- Ability to leverage multiple tools and programming languages to analyze and manipulate data sets from disparate data sources
- Knowledge of API development
- Experience with schema design and dimensional data modeling
- Knowledge of microservices and SOA
- Formal SAFe and/or agile experience.
- Previous healthcare experience and domain knowledge
- Experience designing, building, and maintaining data processing systems
- Experience architecting and building data warehouse and data lakes
Skills
- SQL
- NoSQL
- Python
- Data Warehousing
- ETL/ELT
- Data Pipelines
- Reporting Tools
- Analytic Tools
- Query Optimization
- Data Structures
- Data Transformation
- Metadata Management
- Dependency Management
- Workload Management
- Big Data
- Cloud Architecture
- GCP
- Kubernetes
- AKS
- Kafka
- Spark Streaming
- Requirements Solicitation
- Stakeholder Management
- Complex Systems Analysis
- Analytical Problem Solving
- Collaboration
- Communication
- Data Visualization
- DevOps
- Git
- CI/CD
- Bash Scripting
- UNIX Utilities
- UNIX Commands
- Software Development Methodologies
- Agile
- Waterfall
- API Development
- Schema Design
- Dimensional Data Modeling
- Microservices
- SOA
- SAFe
Location
- NYC
- Hartford, CT
- Wellesley, MA
- Remote
Work Type
- Onsite (2 days/week)
- Full-time Remote
- Full time
Experience Level
- 1+ years of experience
Education Level
- Bachelor’s Degree or equivalent work experience in Computer Science, Information Systems, Data Engineering, Data Analytics, Machine Learning, or related field
- Master’s Degree preferred
Salary/Compensations
- $79,310.00 - $173,040.00
Benefits
- Medical coverage
- Dental coverage
- Vision coverage
- Paid time off
- Retirement savings options
- Wellness programs
- CVS Health bonus, commission or short-term incentive program
About the Company
- We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience.
- At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do.
- Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.
- Our journey calls for technical innovators and data visionaries: come help us pave the way.
- At CVS Health, we possess an extensive repository of healthcare data that spans over 150 million individuals, providing an unparalleled foundation for ambitious Data Engineers.
- As leaders in healthcare, our analytics and engineering teams deliver innovative solutions to business problems.
- Our people fuel our future.
- Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.
Equal Opportunity
- Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.
