About the Role
We are looking for a Data Scientist III to help design, build, and evaluate advanced AI capabilities supporting LeapSpace and Elsevier’s Search & AI Platform initiatives. This role focuses on applied AI development, retrieval systems, and AI evaluation, helping bring cutting-edge AI technologies into production experiences used by researchers worldwide. You will work closely with senior data scientists, engineers, product managers, and domain experts across retrieval systems, generative AI, reasoning workflows, evaluation frameworks, and experimentation, contributing to the next generation of AI-powered scientific discovery tools. This role is ideal for someone with hands-on experience in applied AI, NLP, information retrieval, and LLM-based applications, who enjoys building innovative solutions and translating emerging AI techniques into impactful product capabilities.
Responsibilities
- Develop and improve LLM-powered research workflows, including scientific question answering, literature summarization, semantic exploration and discovery, research insight generation, and citation-aware retrieval and reasoning workflows.
- Build and iterate on agentic and multi-step AI workflows using frameworks such as LangGraph and related orchestration tools.
- Apply modern techniques in NLP, Generative AI, Embeddings and semantic representations, Retrieval-Augmented Generation (RAG), and AI reasoning and workflow orchestration.
- Evaluate emerging AI models, tools, and frameworks and contribute recommendations for experimentation and adoption.
- Contribute to prompt engineering, grounding strategies, context management, and hallucination mitigation efforts.
- Support integration of scientific metadata, ontologies, and knowledge assets into AI-powered workflows.
- Design, develop, and optimize search and retrieval pipelines, including lexical, vector, and hybrid retrieval approaches.
- Contribute to the development and enhancement of RAG systems that integrate LLMs with trusted scientific and biomedical content.
- Experiment with embeddings, re-ranking models, chunking strategies, and retrieval orchestration techniques to improve relevance and answer quality.
- Support development of semantic search, ranking, and knowledge discovery capabilities.
- Collaborate with engineering teams to deploy and scale AI-powered solutions.
- Develop and apply evaluation frameworks for search and AI systems, including IR metrics (e.g., NDCG, recall, precision) and LLM and RAG evaluation metrics (e.g., grounding, faithfulness, hallucination detection).
- Build and maintain evaluation datasets, benchmark suites, and annotation workflows.
- Conduct offline experiments and contribute to online experimentation and A/B testing.
- Analyze experimental results and communicate findings to stakeholders.
- Contribute to responsible AI practices focused on quality, reliability, and trust.
- Partner with product managers, engineers, UX researchers, and domain experts to deliver AI-powered capabilities.
- Communicate technical findings and recommendations clearly to both technical and non-technical audiences.
- Contribute to knowledge sharing and adoption of best practices across the Platform Data Science organization.
- Support delivery of projects from research and experimentation through production deployment.
Requirements
- Master’s or PhD in Computer Science, Data Science, Machine Learning, NLP, Information Retrieval, or a related field
- Experience in data science, machine learning, applied NLP, information retrieval, generative AI, or a related field
- Hands-on experience with LLM-based applications and generative AI systems, RAG pipelines and retrieval systems, search and retrieval architectures (lexical, vector, hybrid), and evaluation methodologies for IR and generative AI systems.
- Strong programming skills in Python
- Experience with modern AI/ML frameworks and tooling (e.g., PyTorch, Hugging Face, LangChain, LangGraph, Haystack)
- Experience working with Databricks or similar distributed data and machine learning platforms
- Understanding of experimentation methodologies, evaluation frameworks, and statistical analysis
- Proficiency with data visualization and analytical tooling (e.g., Tableau, Power BI, matplotlib, seaborn)
- Demonstrated ability to independently execute technical projects and contribute to cross-functional initiatives
Skills
- Applied AI
- NLP
- Information Retrieval
- LLM-based applications
- Generative AI
- Embeddings
- Semantic representations
- Retrieval-augmented generation (RAG)
- AI reasoning
- Workflow orchestration
- Prompt engineering
- Grounding strategies
- Context management
- Hallucination mitigation
- Search and retrieval pipelines
- Lexical retrieval
- Vector retrieval
- Hybrid retrieval
- Embeddings
- Re-ranking models
- Chunking strategies
- Retrieval orchestration
- Semantic search
- Ranking
- Knowledge discovery
- IR metrics
- LLM evaluation metrics
- RAG evaluation metrics
- Evaluation datasets
- Benchmark suites
- Annotation workflows
- Offline experiments
- Online experimentation
- A/B testing
- Responsible AI practices
- Python
- PyTorch
- Hugging Face
- LangChain
- LangGraph
- Haystack
- Databricks
- Data visualization
- Tableau
- Power BI
- matplotlib
- seaborn
Location
- Remote
Work Type
- Full-time
Experience Level
- Data Scientist III
Education Level
- Master’s or PhD in Computer Science, Data Science, Machine Learning, NLP, Information Retrieval, or a related field
Benefits
- Comprehensive Pension Plan
- Home, office, or commuting allowance.
- Generous vacation entitlement and option for sabbatical leave
- Maternity, Paternity, Adoption and Family Care leave
- Flexible working hours
- Personal Choice budget
- Internal communities and networks
- Various employee discounts
- Recruitment introduction reward
- Employee Assistance Program (global)
About the Company
- Elsevier’s mission is to help researchers, clinicians, and life sciences professionals advance discovery and improve health outcomes through trusted content, data, and analytics.
- This role sits within Elsevier’s Platform Data Science organization, a centralized AI and data science group responsible for advancing intelligent discovery, retrieval, and generative AI capabilities across Elsevier products and platforms. The organization develops foundational AI technologies that power experiences such as LeapSpace, Elsevier’s AI-powered research assistant, as well as Elsevier’s broader Search & AI Platform.
- The Platform Data Science organization works at the intersection of: Search and retrieval systems, Generative AI and LLM applications, AI evaluation and experimentation, Semantic enrichment and knowledge systems, Scalable AI platforms and intelligent workflows.
- As a global leader in information and analytics, we help researchers and healthcare professionals advance science and improve health outcomes for the benefit of society. Building on our publishing heritage, we combine quality information and vast data sets with analytics to support visionary science and research, health education, and interactive learning, as well as exceptional healthcare and clinical practice. At Elsevier, your work contributes to the world’s grand challenges and a more sustainable future. We harness innovative technologies to support science and healthcare to partner for a better world.
Equal Opportunity
- We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.
- We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.
- USA Job Seekers: EEO Know Your Rights.
