Research Scientist / Engineer – Performance Optimization
Luma AI · Redwood City
Don’t apply blind. See how your CV matches Engineer first — free, in 30 seconds.
You will leave NewLuxJob. We do not receive or handle applications.
- Company
- Luma AI
- Location
- Redwood City
- Employment type
- Full-time
- Posted
- July 24, 2026
About this job
You'll make Luma's multimodal models fast — profiling and optimizing GPU, CPU, and accelerator code so they train efficiently and deploy at scale without sacrificing quality. You'll write the kernels and operations that get the most out of the hardware. This is deep performance work: fused kernels, tensor cores, Triton and CUDA, distributed multi-node deployment. It fits someone with expert GPU-optimization skills and a deep understanding of transformer internals. If you're not at home in CUDA, Triton, and profilers, this is the wrong depth. What You'll Own - Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency. - Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.…
This is a short summary.
Want to know if you're a fit? Check your CV against this role — free, in 30 seconds.
Most-requested skills in United States
Based on 371 United States vacancies that list requirements, these are the skills employers ask for most often.
- Project Management35% of postings
- CAD24% of postings
- AutoCAD21% of postings
- MATLAB18% of postings
- Revit12% of postings
- Lean11% of postings
- SolidWorks11% of postings
- FEM/FEA7% of postings
Similar jobs
See if your CV fits this job
Paste your CV for an instant match score against this role — and get a tailored cover letter in one click.
- Instant match score for this role
- Tailored cover letter in one click
- Free — no credit card