About the Role
We're hiring a Research Scientist to help build the vision-language models that turn Ironsite's data into the intelligence layer we're building. You'll work alongside our Chief Science Officer and research team on the training, benchmarking, and deployment of state-of-the-art VLMs that can interpret the complexity of a real construction site. Your work will run on data no other lab has, and your models will ship to jobsites where they actually change how things get built.
Responsibilities
- Design, train, and iterate on vision-language models fine-tuned for spatial intelligence in construction environments.
- Run experiments end to end, from data preparation through training, evaluation, and post-training.
- Contribute to the frontier of what's possible on our data, whether that's establishing new baselines, improving post-training recipes, or exploring long-context architectures.
- Contribute to our Construction Intelligence Benchmark suite across video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning.
- Design evaluation metrics that measure real-world construction task performance, not just standard academic benchmarks.
- Own the evaluation loop for your own work so we always know what's actually improving.
- Take your best models from research code to production deployment, working closely with our infrastructure and hardware teams.
- Apply distillation, quantization, and model routing techniques so state-of-the-art understanding runs affordably across our growing fleet.
- Learn from what happens in the field. Field data and production feedback should shape your next experiment.
- Learn from senior researchers, contribute to the intellectual culture of the team, and start to develop your own point of view on what Ironsite's research should look like at scale.
- Read papers, share ideas, and help set the technical bar for the whole team.
Requirements
- 2-4 years of hands-on research experience designing and training deep learning models, particularly transformer-based architectures.
- Industry experience, top-tier PhD program, or equivalent.
- Deep expertise with modern deep learning frameworks (PyTorch, JAX, or similar).
- Strong proficiency in Python with solid software engineering fundamentals.
- Experience working with large-scale vision or language datasets.
- A track record of shipping meaningful research results, whether at a company, in a lab, or in publications.
- A background in Computer Science, Machine Learning, AI, Robotics, or a related field.
- Experience with vision-language models, video understanding, or multimodal architectures.
- Hands-on experience with post-training techniques for large language or vision-language models (SFT, RL methods, parameter-efficient tuning such as LoRA).
- Familiarity with the challenges of video data, including temporal reasoning and long-context modeling.
- Publications at top-tier AI, ML, or CV conferences.
Skills
- Deep learning models
- Transformer-based architectures
- PyTorch
- JAX
- Python
- Software engineering
- Large-scale vision datasets
- Large-scale language datasets
- Vision-language models
- Video understanding
- Multimodal architectures
- Post-training techniques
- SFT
- RL methods
- Parameter-efficient tuning
- LoRA
- Temporal reasoning
- Long-context modeling
Location
- San Francisco Bay Area
Work Type
- On-site
Experience Level
- 2-4 years
Education Level
- PhD program
- Computer Science
- Machine Learning
- AI
- Robotics
- Related field
Salary/Compensations
- $200k-$400k per year
Benefits
- Significant early-stage equity
- Full benefits including health, dental, vision, and 401(k) with 6% match
- Access to dedicated GPU compute resources for research and experimentation
- Daily catered breakfast and lunch
About the Company
- Ironsite is building the intelligence layer for the physical world.
- We design our own wearable hardware, deploy it alongside craft workers, and transform a shift's footage into a next-morning report.
- Our internal team and purpose-built models label the data overnight and deliver actionable insights to superintendents by 5 AM.
- We are accelerating the speed, efficiency, and predictability of construction, especially for complex, mission-critical infrastructure projects, including data centers, LNG facilities, sports stadiums, hospitals, and other large-scale developments, by training AI models on egocentric construction footage and labor productivity data.
- We are built with a pro-worker philosophy at our core: we believe technology should empower the workforce, not replace it.
- We're working to give craft workers and project leaders better visibility into what's happening on-site, while creating a system where the reality of construction and the chaos of each day is finally available to the people running the project.
- Ironsite is deployed across several of the largest active construction projects in the country.
- To date, we've captured more than 100,000 hours of construction footage across seven states, now process thousands of hours of site activity every day, and maintain a worker opt-out rate below two percent.
- This is enabled by a workforce-first architecture that anonymizes devices, captures no audio, and never releases raw video.
- Ironsite is backed by leading investors (8VC, South Park Commons, Saga Ventures) and prominent operators across technology and construction, including Eric Schmidt, Jeff Dean, Jeff Rothschild, Mark Leslie, Scott Wu, Eric Glyman, Karim Atiyeh, Russell Kaplan, and others, alongside over a dozen construction industry operators who have joined us as partners in building this.
- Longer term, we believe Ironsite is the foundation for what construction becomes in the next decade.
- We think the systems we're building are the operating system for how the physical world gets built, and will unlock a fundamentally different way of respect for our workforce.
- One where craft workers are more valued, more visible, and better paid for the skill they bring, and where the industry finally has the intelligence layer that makes autonomous construction possible.
- Both futures start with the same foundation.
