SpaceX
Hawthorne, CA +2
Site Reliability Engineer - AI Infrastructure - Starshield
We tailor your resume to this role and apply for you in seconds.
Apply to Site Reliability Engineer - AI Infrastructure - Starshield at SpaceXJob details
- Location
- Hawthorne, CA +2
- Work type
- Onsite
- Posted
- 3 days ago
- Apply on
- boards.greenhouse.io
About this role
## RESPONSIBILITIES:
- Manage GPU/CPU infrastructure deployments to Top Secret data centers
- Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
- Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
- Develop automation to deploy and manage on-premise Kubernetes\AI clusters, and operating systems
- Deploy and manage core infrastructure such as databases, monitoring and distributed storage
- Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
- Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
- Monitoring and alerting supporting systems to have high availability
- Identify areas for improvement and create innovative solutions that enable high system availability
## BASIC QUALIFICATIONS:
- Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree
- 1+ years of professional experience with Linux operating systems
- Experience with Terraform, Ansible, or other infrastructure tools
- Experience with containerization technologies (i.e. OCI containers, Kubernetes)
- Experience scripting in Bash, Python, or other similar languages
- Development experience in Python, C++, or Go
Ready to apply to SpaceX?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.