Web Crawling - Research Engineer at Thinking Machines Lab | CA, US | Rezi

Web Crawling - Research Engineer at Thinking Machines Lab

Web Crawling - Research Engineer

Thinking Machines Lab · CA, US

Today

Web Crawling - Research Engineer

Thinking Machines Lab · CA, US

8 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Web Crawling - Research Engineer role.

Rezi rewrites your resume against Thinking Machines Lab's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Web Crawling - Research Engineer posting at Thinking Machines Lab — free, in seconds.

About the Role

We are hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep. You will write and own production systems, including the crawler, its infrastructure, and pipelines that turn raw crawls into usable pretraining data. This is fundamentally an engineering role, not a research one.

Responsibilities

  • Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data
  • Build pipelines for large-scale extraction, deduplication, and data quality filtering
  • Build specialized crawlers for high-value or hard-to-reach data sources
  • Work with the pretraining team to understand how changes in crawled data affect model performance
  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale
  • Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned

Requirements

  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
  • A track record of owning crawler or data-acquisition infrastructure at internet scale
  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems
  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team
  • Experience designing systems for petabyte-scale storage and processing
  • Track record of open-source contributions to crawling, scraping, or data infrastructure tools
  • Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team

Skills

  • Python
  • Go
  • Rust
  • Distributed systems
  • Machine learning
  • Web crawling
  • Web scraping
  • Data acquisition
  • Data infrastructure

Location

  • San Francisco, CA

Work Type

  • Onsite

Experience Level

  • 8+ years

Salary/Compensations

  • $350,000-$475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

Equal Opportunity

  • As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
  • Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.