About the Role
Pinterest Labs is seeking research engineers and scientists for the Visual Modeling team to develop and train large-scale visual encoders from scratch. This role involves working with rich visual-text datasets to power visualization features, build multimodal representations for applications like recommender systems, and contribute to the foundation ML models that leverage Pinterest's extensive data.
Responsibilities
- Prototype state-of-the-art visual encoders that power Pinterest's recommender systems and internal visual language models.
- Experiment with billion-scale datasets and gain hands-on experience with large-scale GPU computing.
- Build flexible visual reasoning tools such as composed image retrieval, promptable detection/segmentation, and instruction-tuned embedding and generative models.
- Read research papers, participate in group discussions, and help brainstorm the company's overall visual generative strategy.
- Help collect relevant visual instruction training data that can be shared across multimodal representation, composed image retrieval, text-to-image generation and visual language modeling.
- Publish and share your work through conferences, paper submissions, and blog posts.
- Mentor junior researchers and research interns within the Pinterest Labs organization.
Requirements
- Research engineers and scientists with experience building and training computer vision models.
- Experience with multimodal representations and visual language modeling is strongly preferred.
- A track record of research contributions (e.g., publications, open-source work) and/or shipping ML models to production.
- Hands-on experience with large-scale model training and modern deep learning frameworks (e.g., PyTorch).
- Strong collaboration skills and a demonstrated ability to work effectively in a small, fast-moving team.
- M.S. or PhD in Machine Learning or related academic areas, or equivalent work experience.
- Publications at top ML conferences
- Experience using Cursor, Copilot, Codex, or similar AI coding assistants for development, debugging, testing, and refactoring.
- US based applicants only
Skills
- Computer vision
- Multimodal representations
- Visual language modeling
- Large-scale model training
- Deep learning frameworks (e.g., PyTorch)
- Collaboration
- AI coding assistants (Cursor, Copilot, Codex)
Location
- Remote (situated anywhere in the country)
Work Type
- Hybrid (in-office 1-2 times/quarter for collaboration)
Experience Level
- Research Engineer/Scientist
Education Level
- M.S. or PhD in Machine Learning or related academic areas, or equivalent work experience.
Salary/Compensations
- $189,308—$389,753 USD
Benefits
- Eligible for equity
- Information regarding benefits available for this position can be found here.
About the Company
- Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime.
- At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product.
- Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work.
- At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact.
- Within Pinterest, the Pinterest Labs organization focuses on applied ML research and development.
- Labs works across a broad variety of AI/ML initiatives—including core computer vision, multimodal representation learning, heterogeneous graph neural networks, generative modeling, and recommender systems.
- This is the group that develops the foundation ML models that fully leverage the tens of billions of Pins and the associated knowledge graph to improve the core product.
- We are currently hiring for the Visual Modeling team in Labs, which develops Pinterest's in-house visual encoder.
- The core visual pod is a small group (~10 engineers) inside Labs, which allows for deep collaboration.
- Engineers working on multimodal representation also contribute to our internal text-to-image generation Canvas project—collaborating on autoencoder design or on reward function development for RL training.
Equal Opportunity
- Pinterest is an equal opportunity employer and makes employment decisions on the basis of merit.
- We want to have the best qualified people in every job.
- All qualified applicants will receive consideration for employment without regard to race, color, ancestry, national origin, religion or religious creed, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, age, marital status, status as a protected veteran, physical or mental disability, medical condition, genetic information or characteristics (or those of a family member) or any other consideration made unlawful by applicable federal, state or local laws.
- We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.
- If you require a medical or religious accommodation during the job application process, please complete this form for support.
