Data Scientist - Fintech
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Build and own web-crawling systems that source pretraining data at internet scale, from distributed collection through filtering and deduplication. You'll design the crawler infrastructure, build data pipelines, and work with the pretraining team to understand how crawled data affects model quality. This role requires 8+ years building and scaling web crawlers or large-scale data-acquisition systems, strong software engineering in Python, Go, or Rust with distributed systems experience, and knowledge of web crawling's practical and legal considerations. Based in San Francisco. $350,000–$475,000 annually.
Written from this posting by Neural Jobs AI. The full description is below.
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.
The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.
Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data
Build pipelines for large-scale extraction, deduplication, and data quality filtering
Build specialized crawlers for high-value or hard-to-reach data sources
Work with the pretraining team to understand how changes in crawled data affect model performance
Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale
Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned
8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
A track record of owning crawler or data-acquisition infrastructure at internet scale
Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems
Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)
Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title
Experience designing systems for petabyte-scale storage and processing
Track record of open-source contributions to crawling, scraping, or data infrastructure tools
Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team
Location: This role is based in San Francisco, CA.
Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.
Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.
Here is what this employer asked for. Sign in and we will fill in your half.
Thinking Machines Lab is an AI research and product company developing advanced models and tools that can adapt to individual needs. Its work focuses on making powerful AI more understandable, customizable, multimodal, and useful for collaboration with people. The company combines frontier research with products for model use and customization, including tools that let developers work with open-weight models. Thinking Machines Lab states a broader goal of giving more people access to the knowledge and capabilities required to shape AI for their own applications.
36 more Research Engineer roles in San Francisco
209K – 370K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 37 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Already have an account? Sign in
Continue without an account and apply on the Thinking Machines Lab website
Search by role, company, or anything a posting mentions.