Senior Software Engineer - Reliability Engineering (Remote)
We tailor your resume to this role and apply for you in seconds.
Apply to Senior Software Engineer - Reliability Engineering (Remote) at HomedepotJob details
- Location
- Georgia
- Work type
- Remote
- Compensation
- $90,000 - $170,000/yr
- Posted
- yesterday
- Apply on
- homedepot.wd5.myworkdayjobs.com
About this role
With a career at The Home Depot, you can be yourself and also be part of something bigger.
Position Purpose:
As a Senior Software Reliability Engineer on the Platform Reliability Engineering team, you will ensure the resilience, performance, and security of our enterprise Cloud Platform, working alongside partner RE and Enablement teams supporting our critical Engineering Experience services. As a subject matter expert in Reliability Principles and Practices, you will act as an anchor for our Cloud Platform, mentoring junior engineers — leading incident triage, root cause analysis, driving no-repeat resolutions of systemic problems through ownership of blameless postmortems. Your mission is to ensure that reliability is engineered into our platforms through automation, rigorous change, incident, problem management, and destructive testing, while establishing and enforcing Service Level Objectives (SLOs) that let product teams build and run customer-facing workloads on highly available, paved-path solutions.
Key Responsibilities:
- 50% Delivery and Execution - Develops, tests, deploys, and maintains software, with a clear understanding of the value the software is to provide; Takes on new opportunities and tough challenges with a sense of urgency, high energy and enthusiasm; Consistently achieves results, even under tough circumstances; Develops test suites (functional, destructive, etc) to enable success, rapid deployment of code to production; Takes a broad view when approaching issues; using a global lens
- 20% Learns and Grows - Learns through successful and failed experiment when tackling new problems; Actively seeks ways to grow and be challenged using both formal and informal development channels
- 20% Plans and Aligns - Collaborates with other team members in agile processes; Creates new and better ways for the organization to be successful; Works the Product Team to ensure user stories are valuable, developer ready, easy to understand and testable; Delivers multi-mode communications that convey a clear understanding of the unique needs of different audiences; Adapts approach and demeanor in real time to match the shifting demands of different situations; Relates openly and comfortably with diverse groups of people
- 10% Supports and Enables - Helps grow junior engineers by providing guidance on modern software development frameworks, and leading technical discussions
Direct Manager/Direct Reports:
- This position typically reports to Software Engineer Manager or Sr. Manager
- This position has 0 Direct Reports
Travel Requirements:
- No travel required.
Physical Requirements:
- Most of the time is spent sitting in a comfortable position and there is frequent opportunity to move about. On rare occasions there may be a need to move or lift light articles.
Working Conditions:
- Located in a comfortable indoor area. Any unpleasant conditions would be infrequent and not objectionable.
Minimum Qualifications:
- Must be eighteen years of age or older.
- Must be legally permitted to work in the United States.
Preferred Qualifications:
- 3-5 years of relevant work experience in a related engineering field (Systems, Software, Operational) or a reliability engineering domain
- Deep understanding of and extensive experience with ITIL processes and the support/maintenance of production systems, including Change, Incident, and Problem Management
- Experience leveraging AI tooling to compress MTTD, MTTM, and MTTR
- Extensive experience with common scripting/programming languages (BASH, Python, Golang, Typescript, Java) and data serialization/configuration DSLs (YAML, JSON, HCL)
- Extensive experience with infrastructure automation & orchestration tools such as Terraform and Ansible
- Extensive experience managing Google Cloud Platform (or equivalent) projects and services, including infrastructure, Compute, Developer Tools, Security, and Identity Access Management
- Experience with observability and monitoring tooling such as Prometheus, Grafana, and OpenTelemetry
- Strong understanding of container orchestration (Kubernetes/GKE) and modern microservice architectures
- Familiarity with both Unix/Linux operating systems
- Experience with security tooling (Wiz) & frameworks
- Experience designing and executing destructive, performance, and failure-scenario tests, including leading team drills that validate operational readiness
- Experience with modern debugging and root cause analysis techniques
- Experience with version control systems
- Deep understanding of SLOs and core SRE principles and practices
- Extensive experience taking a lead role in managing live production incidents and problem management, including reporting business impact to leadership
- Strong communication and collaboration skills, with experience producing operational status communications, real-time reporting to diverse stakeholders, documentation, and peer mentorship
Minimum Education:
- The knowledge, skills and abilities typically acquired through the completion of a bachelor's degree program or equivalent degree in a field of study related to the job.
Preferred Education:
- No additional education
Minimum Years of Work Experience:
- 3
Preferred Years of Work Experience:
- No additional years of experience
Minimum Leadership Experience:
- None
Preferred Leadership Experience:
- None
Certifications:
- None
Competencies:
- Global Perspective
- Manages Ambiguity
- Nimble Learning
- Self-Development
- Collaborates
- Cultivates Innovation
- Situational Adaptability
- Communicates Effectively
- Drives Results
- Interpersonal Savvy
For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents: