
Data Scientist, AI Data Foundations
NextDeavor · United States
- Hybrid
- Full-time
- $175,000 / year
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
Job highlights
- Build AI/ML data foundations for training.
- Design and operate vector, feature, graph stores.
- Analyze data to find actionable insights.
- Engineer curated AI-ready datasets.
- Collaborate with ML engineers and scientists.
About the role
Data Scientist, AI Data Foundations
Become a Key Player as a Data Scientist, AI Data Foundations. You will design and build the curated data structures that AI and ML applications consume, enabling higher-quality model training and inference. You will partner with model builders, product, risk, and growth stakeholders to surface actionable insights and ship production-ready vector, feature, and graph data assets. This is a Remote role.
How You'll Make an Impact on the Team
- Build and maintain vector stores for RAG, including embedding pipelines, chunking strategies, indexing, and refresh patterns.
- Own the feature store: design, build, and operate feature definitions, freshness SLAs, lineage, and point-in-time correctness for offline/online use.
- Design and implement graph data structures to model relationships across applicants, applications, products, lenders, decisions, and outcomes.
- Lead data discovery: profile lending, deposit, and behavioral datasets to identify trends, segments, anomalies, and model drivers; produce actionable hypotheses for stakeholders.
- Engineer curated, AI-ready datasets with appropriate quality checks, documentation, and governance for downstream model builders and analysts.
- Define and run evaluation frameworks for RAG retrieval quality, feature drift, embedding quality, and graph completeness; iterate on metrics.
- Partner closely with ML engineers and applied scientists to ensure data assets accelerate model development and serving workflows.
- Champion responsible data use by collaborating with governance, security, and compliance teams to ensure data classification, consent, and regulatory boundaries are respected.
- Communicate findings via write-ups, notebooks, dashboards, and short presentations for technical and non-technical audiences.
What You'll Need to Be Successful in This Role
- 4–7 years of experience in data science, ML engineering, or applied data roles, with significant time building data assets consumed by models or applications.
- Hands-on experience designing and operating vector stores for RAG or semantic search (embedding generation, chunking, indexing, retrieval evaluation).
- Experience building or operating a feature store (e.g., Databricks Feature Store, Feast, or custom), including offline training and online serving patterns and point-in-time correctness.
- Experience modeling and building graph data structures and writing graph queries (Neo4j, TigerGraph, Cosmos DB Gremlin, or similar).
- Strong proficiency in Python (pandas, NumPy, scikit-learn, PySpark) and SQL; comfortable using Databricks notebooks and jobs.
- Practical experience with embedding models and LLM tooling (Hugging Face, OpenAI/Azure OpenAI APIs, LangChain or similar) in production or near-production contexts.
- Demonstrated data discovery skills: profiling messy datasets, surfacing patterns, validating findings statistically, and explaining results clearly.
- Solid grounding in classical ML concepts (supervised vs. unsupervised learning, train/test discipline, leakage, evaluation metrics).
- Strong written and verbal communication skills for technical and business audiences.
What Else Might Help You Out
- Experience in SaaS or FinTech, especially with lending, deposit, credit, fraud, or KYC/AML data.
- Familiarity with Databricks-native AI/ML tooling: Databricks Vector Search, Databricks Feature Store, MLflow, Unity Catalog.
- Experience with open-source vector DBs (pgvector, Pinecone, Weaviate, Chroma, FAISS) and strong opinions on trade-offs.
- Experience with Microsoft Azure data and AI services (Azure OpenAI, Azure AI Search, ADLS Gen2).
- Experience evaluating RAG systems end-to-end (recall@k, faithfulness, answer quality, hallucination measurement).
- Exposure to graph algorithms (community detection, link prediction, centrality) applied to business problems.
- Bachelor's or Master's in CS, Statistics, Mathematics, Engineering, or related quantitative field, or equivalent experience.
Pay Range
$114,000 - $175,000/year
Ready to Make Your Mark?
This role may fill quickly. Submit your resume to be considered.
Key skills/competency
- Data Scientist
- AI Data Foundations
- Vector Stores
- Feature Stores
- Graph Data Structures
- Data Discovery
- ML Engineering
- Python
- SQL
- LLM Tooling
Skills & topics
- Data Scientist
- AI Data Foundations
- Machine Learning
- Vector Stores
- Feature Stores
- Graph Databases
- Python
- SQL
- LLM
- RAG
- Data Engineering
- Remote
- NextDeavor
How to get hired
- Tailor your resume: Highlight experience with vector stores, feature stores, graph data structures, and Python/SQL.
- Showcase relevant projects: Detail your work with LLM tooling, RAG systems, and data discovery.
- Quantify achievements: Use numbers to demonstrate the impact of your data engineering and ML contributions.
- Prepare for technical questions: Be ready to discuss ML concepts, data modeling, and your experience with specific tools.
- Demonstrate communication skills: Practice explaining complex technical findings to both technical and non-technical audiences.
Technical preparation
Master Python for data manipulation and ML.,Practice SQL for complex data querying.,Build projects using vector and feature stores.,Familiarize with LLM and RAG concepts.
Behavioral questions
Describe a complex data problem you solved.,How do you ensure data quality and governance?,Explain a technical concept to a non-technical person.,How do you collaborate with cross-functional teams?
Frequently asked questions
- What is the primary focus of the Data Scientist, AI Data Foundations role at NextDeavor?
- The Data Scientist, AI Data Foundations role at NextDeavor focuses on designing and building curated data structures for AI and ML applications, specifically including vector stores for RAG, feature stores, and graph data structures. The goal is to enable higher-quality model training and inference.
- Is this Data Scientist position remote?
- Yes, this Data Scientist, AI Data Foundations position at NextDeavor is a fully remote role, allowing you to work from anywhere.
- What kind of data assets will I be working with as a Data Scientist at NextDeavor?
- As a Data Scientist, AI Data Foundations at NextDeavor, you will work with vector, feature, and graph data assets. This includes data related to lending, deposits, and user behavior, engineered into AI-ready formats.
- What technical skills are most important for the Data Scientist, AI Data Foundations role?
- Key technical skills for this role include hands-on experience with vector stores (RAG, embedding, indexing), feature stores, graph data modeling, Python (pandas, NumPy, scikit-learn, PySpark), SQL, and LLM tooling (Hugging Face, OpenAI APIs, LangChain).
- What experience level is NextDeavor looking for in a Data Scientist?
- NextDeavor is looking for candidates with 4-7 years of experience in data science, ML engineering, or applied data roles, specifically with experience in building data assets consumed by models or applications.
- How can I highlight my qualifications for the Data Scientist role on my resume?
- To highlight your qualifications for the Data Scientist, AI Data Foundations role, focus on quantifiable achievements in building vector stores, feature stores, and graph data structures. Emphasize your proficiency in Python, SQL, and LLM tooling, as well as your experience in data discovery and ML concepts.
- What is the salary range for the Data Scientist position at NextDeavor?
- The salary range for the Data Scientist, AI Data Foundations position at NextDeavor is $114,000 to $175,000 per year.
- What are the potential career growth opportunities for a Data Scientist at NextDeavor?
- While specific growth paths aren't detailed, this role offers opportunities to work with cutting-edge AI and ML technologies, build foundational data assets, and partner with various stakeholders, which can lead to advancements in AI/ML engineering, data science leadership, or specialized AI roles within the company.