About the Role
As a Data Scientist, you will support the development, evaluation, and deployment of AI and NLP models used in our products. You will contribute to building and optimizing components of our GenAI, RAG, and NLP pipelines, working closely with senior data scientists, software engineers, and subject matter experts. This role is ideal for candidates with foundational knowledge of NLP/ML who are eager to learn, take ownership of tasks, and grow into more advanced responsibilities.
Responsibilities
- Assist in collecting, cleaning, and preparing structured and unstructured data for model development.
- Contribute to ML/NLP model prototyping and evaluation, including defining basic quality metrics with guidance.
- Support optimization of Retrieval Augmented Generation (RAG) components such as document ingestion, retrieval, and preprocessing.
- Help design and test model inference pipelines, transformer-based models, and GenAI applications.
- Develop well-structured, production-ready Python modules for preprocessing, model execution, and evaluation.
- Collaborate with developers to integrate data science components into production workflows.
- Support end-to-end quality validation of deployed pipelines and help monitor model performance over time.
- Stay informed about emerging ML and NLP techniques and tools relevant to ongoing projects.
- Communicate results in a clear, structured way to technical and non-technical stakeholders.
- Work with senior team members on both independent tasks and small-scale project ownership.
Requirements
- MSc/MTech in Computer Science, Data Science, Artificial Intelligence, Mathematics, Statistics, Bioinformatics, or a similar quantitative field OR Bachelor’s degree + 1–2 years of relevant experience.
- Some applied experience with ML or NLP projects (internships, academic work, or industry roles).
- International study or work exposure is a plus.
- Solid Python programming skills for data analysis and model development.
- Basic familiarity with transformer models and modern NLP techniques.
- Experience using LLMs via APIs, prompt engineering, or experimentation with GenAI tools.
- Understanding of RAG concepts and willingness to learn how to implement them.
- Exposure to cloud environments (AWS, Azure, Bedrock) or interest in learning deployment workflows.
- Familiarity with common ML algorithms (e.g., logistic regression, SVMs, random forests) and model evaluation practices.
- Experience with GitHub/GitLab and Agile ways of working.
- Strong analytical and problem-solving mindset.
- Curiosity and willingness to learn new technologies.
- Clear communication and ability to work in a collaborative, cross-functional team.
Skills
- GenAI
- ML
- NLP
- Python
- transformer models
- LLMs
- RAG
- AWS
- Azure
- Bedrock
- logistic regression
- SVMs
- random forests
- GitHub
- GitLab
- agent frameworks
- LangChain
- Databricks
- OpenSearch
- deep learning
- neural networks
- transfer learning
- CI/CD workflows
Location
- Amsterdam (Radarweg)
Work Type
- Full-time
Experience Level
- 1-2 years of relevant experience
Education Level
- MSc/MTech in Computer Science, Data Science, Artificial Intelligence, Mathematics, Statistics, Bioinformatics, or a similar quantitative field
- Bachelor’s degree
Salary/Compensations
- €49,000 - €81,700
Benefits
- Country specific benefits
About the Company
- Data Science Corporate Markets is a diverse team focusing on GenAI, ML, NLP. We mainly develop best-in-class enrichment pipelines for Elsevier’s corporate markets .com products such as Reaxys, Embase and Pharmapendium.
Equal Opportunity
- We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.
- EEO Know Your Rights.
