
Developer
NLB Services · United States
- On site
- Full-time
- United States
Tailored resume — keyword-matched to this role.
Hiring manager — we find who's hiring.
Intro email — drafted to reach them directly.
About the role
Job Title: GPU Kernel Developer – AI/ML (Remote)
Location: Remote
Fulltime
About the Role
Wipro is seeking a highly skilled GPU Kernel Developer to join our cutting-edge AI/ML engineering team. This role focuses on optimizing and scaling large language models (LLMs) such as BERT, with emphasis on high-performance tuning and GEMM operations. You will work on developing and optimizing kernels for PyTorch and Triton, ensuring maximum efficiency across GPU architectures.
Key Responsibilities
- Design, implement, and optimize GPU kernels for AI/ML workloads.
- Focus on matrix multiplication (GEMM) and other performance-critical operations.
- Collaborate with researchers and engineers to accelerate BERT and transformer-based models.
- Integrate optimized kernels into PyTorch and Triton frameworks.
- Conduct performance profiling, benchmarking, and tuning across diverse hardware.
- Work closely with system engineers to ensure seamless deployment in production environments.
Must-Have Skills
- Strong expertise in GPU kernel development.
- Hands-on experience with PyTorch and Triton (mandatory).
- Proficiency in Python, C, and C++.
- Deep understanding of system-level programming and GPU architecture.
- Experience with AI/ML model optimization, especially BERT.
- Strong knowledge of high-performance computing (HPC) and parallel programming.
Preferred Qualifications
- Prior experience in large-scale AI/ML model deployment.
- Familiarity with CUDA or other GPU programming frameworks.
- Background in distributed training systems.
- Strong problem-solving and debugging skill