Yotta Labs
Research Engineer Intern - AI Systems
We tailor your resume to this role and apply for you in seconds.
Apply to Research Engineer Intern - AI Systems at Yotta LabsJob details
- Work type
- Remote
- Visa
- Sponsorship available
- Posted
- yesterday
- Apply on
- jobs.ashbyhq.com
About this role
Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization, owning a well-scoped project that impacts AI applications deployed on their platform.
What you'll do:
- Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium
- Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA
- Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations
- Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups
- Ship code upstream to open-source AI infrastructure projects, with tests and documentation
What they're looking for:
- Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field
- Solid programming skills in Python and familiarity with C++
- Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects
- Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count
- Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler)
- Strong problem-solving skills and the ability to work independently in a collaborative, remote environment
- Contributions to open-source AI infra projects like vLLM, SGLang, PyTorch, or Triton
- Familiarity with LLM inference internals — FlashAttention, PagedAttention, continuous batching, speculative decoding, MoE, or quantization
- Experience with profiling tools (e.g. Nsight, ROCm Profiler, Neuron Profiler, or PyTorch Profiler) and performance debugging on real workloads
- Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA
Benefits:
- Flexible remote work environment
- Direct mentorship from engineers from leading institutions and tech companies
- Access to serious hardware — latest-generation NVIDIA GPUs, AMD accelerators, and AWS Trainium at scale
- A fast path to a full-time return offer for top performers
Ready to apply to Yotta Labs?
We tailor your resume to this role and apply for you.