Research Engineer, Preference Data
Vizcom · San Francisco
- Location
- San Francisco
- Posted
- Yesterday
- Type
- Fulltime / Remote
- Salary
- Not listed — worth asking early.
What the job really is
ML Data Engineer
San Francisco, CA · In Person · Full-Time
Applying to this role will also allow us to consider you for other research opportunities at Vizcom. We believe the best roles are shaped around exceptional people, not just job descriptions.
About Vizcom
Vizcom is where design teams at companies like Nike, GM, New Balance, and Hasbro bring ideas from sketch to product. Designers use Vizcom to sketch, render, explore color and materials, work in 3D, and prepare concepts for production.
The render itself was never the point. The point is the physical thing that comes after it. We call this pencil to product.
Vizcom is a Series B company with more than $52M raised.
More than 700,000 designers have worked in Vizcom, and every session leaves a trail: candidates selected, outputs promoted into designs, regions masked and renamed, and entire directions kept or discarded.
That trail is one of the most valuable things we create outside of the product itself. Today, though, it's more archaeology than asset. Only a fraction of what happens in a session reaches training-grade quality, while increasingly sophisticated post-training methods depend on exactly this kind of high-quality, domain-specific data.
Your job will be to turn that trail into a machine.
A design session isn't a simple sequence — it's a branching tree. Designers fork, backtrack, iterate, and abandon entire directions on their way to the thing they ultimately keep. The judgment lives in the shape of that process, and today we capture only pieces of it.
The Role
As an ML Data Engineer, you'll build the data flywheel itself: the systems that turn professional design work into training-grade preference data and training results back into a better product.
You'll work alongside the researchers consuming what you build, within the product systems where these signals originate, and across the data infrastructure where they ultimately land. Your closest users are the researchers sitting beside you, and you'll see quickly when a dataset you've built allows them to ask a question they couldn't ask before.
This is not a support role, and it isn't traditional offline ETL.
The pipelines you design will run through a live product used every day by professional design teams, including enterprise customers with rigorous expectations around privacy and data protection. Capturing better signals without compromising user trust, contractual obligations, or product performance is a core part of the work.
If you want to train models without building the systems that feed them, this probably isn't the role for you. If you believe the next advances in ML will increasingly be won through better data, it might be.
We think about a dataset as a product: it has users, versions, provenance, and a quality bar. A training result should be reproducible from a dataset fingerprint months later, and "Where did this example come from?" should always have an answer.
Here, building that standard is the job.
What You'll Own
The capture surface: Determine what the product records, working closely with Product and Engineering to design and ship instrumentation in production systems.
The data pipeline: Build the path from canvas to warehouse to training set, ensuring data is clean, versioned, reliable, and reproducible.
Dataset contracts and lineage: Make every example traceable to its origin and every training dataset reproducible months later.
Privacy and data boundaries: Build systems that reflect what can be captured and used under different enterprise agreements, with consent, isolation, and appropriate safeguards designed in from the beginning.
Research-ready datasets: Create appropriately governed and sanitized datasets that allow researchers to experiment safely and effectively.
Collection instruments: When historical signals aren't enough, build mechanisms for gathering explicit feedback that designers actually want to use.
Honest representations of ambiguity: Design data systems that preserve context. Unpicked doesn't necessarily mean disliked, abandoned doesn't necessarily mean rejected, and our data should reflect the difference.
This is a charter, not a week-one checklist. We don't expect one person to tackle everything at once. Part of the role is helping determine what matters most and in what sequence.
What Your First 90 Days Could Look Like
Days 1–30: Map
Understand the event surface, warehouse, existing datasets, research workflows, and current data boundaries. Identify what exists, what's missing, and which research questions our data can't yet answer.
Days 30–60: Ship
Take one new signal end to end: instrumented in the product, landed reliably in the warehouse, versioned appropriately, and available to researchers.
Days 60–90: Establish the standard
Define the data standards that future work will build on, including versioning, dataset fingerprints, lineage, quality, and data boundaries.
What We're Looking For
Experience building training-data infrastructure, ML data systems, or large-scale data pipelines that other teams depended on.
Strong software engineering skills and experience building reliable production systems.
Experience designing data models and pipelines with reproducibility, observability, and lineage in mind.
An experimental mindset and comfort working closely with researchers to turn ambiguous questions into measurable datasets.
Strong judgment around data quality, including an understanding of when the absence of a signal is meaningfully different from a negative signal.
Comfort operating in an environment where the underlying systems and standards are still being built.
Nice to Have
You've trained models yourself and understand what ML training pipelines actually need from their data.
You've worked with preference data, labeling systems, human-feedback pipelines, or evaluation operations.
You've built data infrastructure under meaningful privacy, security, or contractual constraints.
You've worked with large-scale event or behavioral datasets.
You have experience with data systems supporting generative AI or multimodal models.
Above all, we're looking for someone who can look at the exhaust of a complex product and see evidence.
What You'll Get
A unique dataset: The recorded decisions and workflows of more than 700,000 designers, with new signals generated every day.
Researchers as your users: You'll work directly alongside the people using the datasets you build, creating an unusually tight feedback loop between data engineering and research.
High ownership: You'll help establish the standards for how Vizcom captures, versions, governs, and uses ML training data.
A rare problem space: High-quality professional preference data is difficult to create. You'll have the opportunity to build systems around a dataset and domain that few ML teams have access to.
Direct access to the founders: You'll work closely with Vizcom's founders and technical leadership as we build out our research and ML infrastructure.
Benefits at Vizcom
100% employer-sponsored medical coverage for employees, plus 25% coverage toward dependents
Dental and vision coverage, plus mental health benefits
Meaningful equity ownership
Flexible PTO
401(k) with employer match
Generous annual Learning & Development allowance
Paid parental leave
Weekly catered lunch at our San Francisco headquarters
Monthly gym membership stipend
Compensation
Base salary: $220,000–$300,000 USD + equity
We regularly benchmark compensation against relevant peer companies using current market data from industry-standard sources, including Carta and Pave. This range reflects our Tier 1 compensation market, which includes San Francisco.
The actual offer and overall compensation package will be determined based on multiple factors, including relevant experience, skills, qualifications, and business considerations. The compensation and benefits described in this posting apply to U.S.-based W-2 employees and may vary based on applicable employment laws and requirements.
How We Work
We document what we learned, not just what we worked on.
Negative results are valuable when they help us close off the wrong paths.
Results should be reproducible before they earn additional compute.
We share meaningful research through technical write-ups, demonstrations, and showcases where appropriate.
Our interview process emphasizes real-world problem solving and practical technical work rather than LeetCode-style interviews.
Location
This is an in-person role based in San Francisco, CA.
Join Us
At Vizcom, we move quickly, give people meaningful ownership, and offer the opportunity to shape both our product and our company as we grow. We believe deeply in the craft of industrial design and in building tools that help designers bring better ideas into the physical world.
Join us in shaping a world designed by you.
A Note to Candidates and Recruiting Agencies
Please apply directly through this job posting. To help us keep our hiring process fair and organized, we ask candidates not to contact Vizcom employees directly regarding their application or candidacy.
Vizcom is not seeking assistance from external recruiting agencies or search firms for this role. Please do not contact or solicit Vizcom employees regarding recruiting services, candidate submissions, or agency partnerships. Unsolicited resumes or candidate profiles submitted by agencies will not create a fee obligation on behalf of Vizcom.
As part of Vizcom's SOC 2 Type II compliance program, employment is contingent upon successful completion of a background check, as permitted by applicable law.
Similar roles
Open positions we recommend based on this role.