Data Scientist (3-5 years)
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
LanceDB
NewBuild end-to-end AI training workflows using LanceDB across different domains, demonstrating how the platform accelerates research from data curation to modeling. You'll need 5+ years training deep learning models, ideally with video or world models, proven success shipping state-of-the-art models, and experience maintaining popular open-source projects. The role involves benchmarking experiments, publishing research, and partnering with engineering and product teams. PyTorch and distributed training experience is valued.
Written from this posting by Neural Jobs AI. The full description is below.
AI advances at the speed of its research, and research moves at the speed of its data. LanceDB is the AI-native Multimodal Lakehouse: one system where a researcher curates petabytes of video, audio, and every signal derived from them with a few lines of Python, and the next training run starts as fast as the next idea. Customers like Runway, Midjourney, and Netflix build the future of AI on LanceDB, from frontier and world models to robots and autonomous vehicles.
As the AI research engineer at LanceDB, you'll work with the research team to perform fairly open ended research, focused on end to end training flows across different AI domains, showcasing how LanceDB can be used to accelerate research flows.
This is an opportunity to pursue your research interest as an engineer, and have a meaningful impact on raising awareness and significantly improve the product.
Show how Lancedb can be used for training models end to end from curation to modeling across industry verticals
Compare the Lancedb stacked workflow with existing standard training flows with well designed and replicable experiments, that may include benchmarking
Provide core content and work cross-function with to increase awareness for workflow specific features like blobv2, distributed indexing etc.
Publish models and research papers on LanceDB blog platform, social media, and in AI conferences
Partner closely with engineering and product to provide feedback from a researcher’s perspective
5+ years of experience in training deep learning models, not limited to LLM, ideally have worked with video, action, world models before
Proven track record of training SOTA models in an industry vertical
Strong experience in building and maintaining popular OSS repos.
Demonstrated ability to map user feedback from noise to key deliverables
Excellent prioritization skills and demonstrate execution efficiency
Strong sense of product GTM, demonstrate ability to balance strategic thinking with hands-on execution
Passion for staying up-to-date with SOTA AI research and trends
Experience with training transformer based models, and post-training/alignment
Hands on experience with PyTorch, distributed training, and tensor parallelism
5+ years of experience, including working at startups
Here is what this employer asked for. Sign in and we will fill in your half.
LanceDB develops a multimodal lakehouse for AI data. Its open-format architecture stores and queries text, images, video, embeddings, and other large datasets together in customer-controlled object storage. AI teams use LanceDB for large-scale data curation, feature engineering, training-data workflows, retrieval, and search. The platform is designed for datasets that are too large or varied for a collection of disconnected tools, combining the Lance format, database capabilities, and cloud services in one data foundation.
Search by role, company, or anything a posting mentions.