About the Role
Join an exciting opportunity focused on advancing next-generation AI technologies. This role is ideal for candidates with strong expertise in Machine Learning, Large Language Models (LLMs), experimental research, and data analysis who are passionate about solving complex challenges in a collaborative, cutting-edge environment.
Responsibilities
- Design, develop, and execute AI/ML experiments and evaluation frameworks.
- Build and improve AI benchmarking methodologies for large language models and related systems.
- Analyze experimental results and provide data-driven insights.
- Develop scalable research tools using Python.
- Collaborate with cross-functional research and engineering teams.
- Maintain clean, reproducible code using Git and modern development practices.
- Contribute to research documentation and technical reports.
Requirements
- Master's or PhD in Computer Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Physics, Computational Sciences, or a related technical field.
- Experience in one or more of the following roles: Research Engineer, Research Scientist, Applied Scientist, Machine Learning Engineer, AI Researcher, Data Scientist.
- Strong proficiency in Python.
- Hands-on experience with Machine Learning, Large Language Models (LLMs), experimentation, and data analysis.
- Experience using Git, IDEs, and Jupyter Notebook/Google Colab environments.
- Excellent analytical, research, and problem-solving skills.
- Experience with AI evaluation, benchmarking, testing, or red teaming.
- Research publications or notable contributions to AI/ML projects.
- Experience designing and conducting experimental research.
- Familiarity with modern AI evaluation methodologies and benchmark development.
- Strong research aptitude and scientific thinking.
- Excellent coding and debugging skills.
- Ability to work independently in a remote environment.
- Passion for advancing state-of-the-art AI systems.
- Strong communication and collaboration skills.
Skills
- Python
- Machine Learning
- Large Language Models (LLMs)
- Data Analysis
- Git
- AI Evaluation
- Benchmarking
- Testing
- Red Teaming
- Experimental Research
- Coding
- Debugging
Location
- Remote (USA)
Work Type
- Remote
- Full-time
Education Level
- Master's Degree
- PhD
