Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius · Palo Alto, California, United States
Posted about 2 months ago
or apply directly on Nebius's site. We never take the application ourselves.
Is this posting real?
- This role has been open
- 66 days Nebius's roles stay open a median of 66 days
- Reposted
- No
- Salary listed
- No 9% of Nebius's roles list one
- Ghost-job risk at Nebius
- high 318 stale, 17 reposted of 384 open
- Hiring momentum
- 508 roles opened in the last 90 days ↑ up vs. the prior 90 days
- Last confirmed on the employer's board
- 2026-10-08
Measured from postings appearing on and disappearing from Nebius's own greenhouse board since 2026-08-03. Full hiring picture for Nebius.
About this role
As a Senior Applied Scientist at Nebius, you will lead research projects focused on optimizing LLM and VLM inference, developing methods for quantization and model compression, and building prototypes using tools like PyTorch and Triton. Your work will involve rigorous experimentation, collaboration with engineering teams, and sharing findings through technical documentation and publications, all aimed at enhancing production capabilities in AI systems.
- benefits
- 3/5
- freshness
- 1/5
- career value
- 5/5
- role clarity
- 5/5
- pay transparency
- 0/5
Scored from the posting itself — how clearly the role is described, how much it says about pay and benefits, and how recently it was listed. Not a judgement of Nebius as an employer.
What you need
- A PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied mathematics, or a closely related discipline.
- A strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
- Strong Python and PyTorch implementation skills, with the ability to turn ideas into experiments and working prototypes.
- Deep knowledge of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving trade-offs.
- Strong experimental design skills covering ablations, baselines, metrics, statistical reasoning, and failure analysis.
- Excellent written and verbal communication.
Nice to have
- First-author publications at venues such as NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, or HPCA.
- Experience deploying ML models or inference optimizations in production.
- Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
- Experience applying post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation to inference quality or efficiency.
- Open-source research artifacts, widely used benchmarks, technical blogs, or invited talks demonstrating contributions to efficient AI systems.
What you get
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Worth weighing
- No specific mention of the team size or structure, which may affect collaboration dynamics.
- The role involves a significant amount of research and experimentation, which may not appeal to those looking for more straightforward engineering tasks.
- The posting emphasizes a fast-moving and bold environment, which may imply high expectations and pressure.
- No specific mention of work-life balance, which could vary depending on project demands.
Summarised from Nebius's posting. Read the full original.
Listed by Nebius on their greenhouse job board, last confirmed open on 2026-10-08. PitchMeAI is not the employer.
More roles at Nebius
- Offensive Security LeadRemote - Europe
- Vulnerability Operation Center LeadRemote - Europe
- Technical Program Manager – Data Center Infrastructure DeploymentsHyderabad, India; India
- Senior Product Manager, Agentic SearchIsrael
- Cloud Solution Architect - Educational Content Author, Nebius AcademyRemote - Europe
- Principal, EMEA GTM - Physical AIAmsterdam, Netherlands; Germany; London, United Kingdom
- Forward Deployed Engineer - Physical AI Cloud PlatformRemote - United States
- Site Reliability Engineer in Hardware InfrastructureAmsterdam, Netherlands