Yotta Labs

Research Engineer Intern - AI Systems

RemotePosted yesterdayVisa Sponsorship

We tailor your resume to this role and apply for you in seconds.

Apply to Research Engineer Intern - AI Systems at Yotta Labs

Job details

Work type
Remote
Visa
Sponsorship available
Posted
yesterday
Apply on
jobs.ashbyhq.com

About this role

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization, owning a well-scoped project that impacts AI applications deployed on their platform.

What you'll do:

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium
  • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA
  • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations
  • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups
  • Ship code upstream to open-source AI infrastructure projects, with tests and documentation

What they're looking for:

  • Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field
  • Solid programming skills in Python and familiarity with C++
  • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects
  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count
  • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler)
  • Strong problem-solving skills and the ability to work independently in a collaborative, remote environment
  • Contributions to open-source AI infra projects like vLLM, SGLang, PyTorch, or Triton
  • Familiarity with LLM inference internals — FlashAttention, PagedAttention, continuous batching, speculative decoding, MoE, or quantization
  • Experience with profiling tools (e.g. Nsight, ROCm Profiler, Neuron Profiler, or PyTorch Profiler) and performance debugging on real workloads
  • Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA

Benefits:

  • Flexible remote work environment
  • Direct mentorship from engineers from leading institutions and tech companies
  • Access to serious hardware — latest-generation NVIDIA GPUs, AMD accelerators, and AWS Trainium at scale
  • A fast path to a full-time return offer for top performers
Ready to apply to Yotta Labs?
We tailor your resume to this role and apply for you.

About Yotta Labs

Yotta Labs