Clera
San Mateo, California

ML Infrastructure Engineer

OnsitePosted 3 days ago
24,178applications sent for our users this week
This job. More like it. We find them and apply for you.
ApplyBolt finds jobs that match you, tailors your resume to each role, and submits applications on company sites. All handled for you.
Find & apply for me

We tailor your resume to this role and apply for you in seconds.

Apply to ML Infrastructure Engineer at Clera

Job details

Location
San Mateo, California
Work type
Onsite
Posted
3 days ago
Apply on
jobs.ashbyhq.com

About this role

About the Role

This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context and data governance layer for AI agents in highly regulated industries. You will own the inference and model-serving infrastructure end to end, ensuring AI agents run reliably, accurately, and at scale in production environments where performance is non-negotiable.

What You'll Do

  • Design, build, and operate inference and model-serving infrastructure from development through production deployment.

  • Scale systems to support AI agents running reliably under increasing concurrency and production load.

  • Identify and resolve infrastructure bottlenecks in close collaboration with ML and platform engineering teams.

  • Optimize systems for latency, throughput, and reliability at scale.

What We're Looking For

  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

  • Hands-on experience designing and scaling inference serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

  • Strong systems engineering fundamentals with expertise in distributed systems, containerization, and orchestration (Docker, Kubernetes).

  • Demonstrated ability to optimize production ML systems for latency, throughput, and reliability under high concurrency.

  • Experience with cloud infrastructure platforms such as AWS, GCP, or Azure for deploying and managing ML workloads.

  • Proficiency with monitoring, observability, and debugging tools such as Prometheus, Grafana, ELK, or distributed tracing frameworks.

  • Proficiency in at least one systems programming or backend language: Python, Go, Rust, C++, or Java.

  • Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.

  • Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines is a plus.

  • Experience with enterprise data infrastructure, data pipelines, or data integration platforms is a plus.

Location

This role is on-site in San Mateo, California. Visa sponsorship is not available.

Ready to apply to Clera?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.

About Clera

Clera
San Mateo, California