Mirrai Careers
Resume BuilderCareer Test
JobsInsightsPricing
Get Started Free
Jobs/Member of Technical Staff - Infrastructure

Member of Technical Staff - Infrastructure

Gimlet Labs

San Francisco, CA Full-time$150k–$350k / year Posted 30+ days ago
Above market. This role pays above the $200k median for similar USD roles (110 comparable postings in our corpus).
Apply on company site
About Us Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together. Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization. We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI. This gives our team access to systems research problems grounded in frontier models, cutting-edge production workloads, and emerging hardware architectures. ABOUT THIS ROLE We are looking for an Infrastructure Platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet's heterogeneous AI cloud. In this role, you will build the platform that brings new hardware online, provisions clusters, manages capacity, and keeps production inference systems running reliably at scale. You'll work across bare metal, Linux, Kubernetes and cluster schedulers, high-speed networking, observability, and automation to ensure AI workloads can execute efficiently in production. Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple accelerator vendors and architectures. You'll build the operational systems that abstract this complexity, allowing new silicon to become production-ready quickly while ensuring workloads remain reliable, observable, and performant from day one. This is a highly hands-on systems role. You'll partner closely with distributed systems, runtime, compiler, networking, and hardware engineers to build the infrastructure foundation that powers the next generation of AI workloads. WHAT SUCCESS LOOKS LIKE In your first 12–18 months, you will help: * Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters powering production AI inference. * Build provisioning and lifecycle management systems that automate deployment, upgrades, validation, and fleet operations. * Improve cluster scheduling, resource utilization, isolation, and capacity management across heterogeneous hardware. * Build highly observable infrastructure that enables rapid debugging, incident response, and operational excellence. * Partner with distributed systems, runtime, compiler, networking, and hardware engineers to bring new accelerator platforms into production. * Influence the architecture of the infrastructure platform that will power the next generation of AI workloads. YOU MAY BE A GOOD FIT IF * Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems. * Deep Linux systems experience, including debugging performance, networking, storage, processes, and kernel-level issues. * Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems. * Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent. * Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation. * Familiarity with high-performance networking such as InfiniBand, RoCE, high-speed Ethernet, or datacenter fabrics. * Strong operational judgment: you know how to build systems that are observable, recoverable, and boring in production. * Comfort working in a fast-moving startup environment with high ownership and ambiguity. * Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience. STRONG CANDIDATES MAY ALSO HAVE * Experience building or operating AI inference, training, HPC, or neocloud infrastructure. * Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up. * Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting. * Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks. * Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling. * Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators. Why join now? Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company. As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company. We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years. Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

See how well you match this job

Upload your resume and we’ll score your fit for this role and 6 similar roles — then tailor your CV to it with AI. Free, no credit card.

Check your match

Similar jobs

  • Member of Technical Staff, Infrastructure

    mandolin

    San Francisco$160k–$270k
  • Member of Technical Staff, Infrastructure

    Vapi

    San Francisco$200k–$280k
  • Network Engineer

    gimlet

    San Francisco, CA$250k–$320k
  • Software Engineer, Compute Infrastructure

    OpenAI

    Remote$230k–$405k
  • Senior Infrastructure Engineer

    Bland AI

    San Francisco$120k–$200k
  • Staff Software Engineer, Inference Infrastructure

    Cohere

    Remote
Apply on company site

Want more roles like this? Browse fresh jobs or tailor your resume with AI.

Mirrai Careers

AI-powered career platform: build resumes, match jobs, and plan your career.

Product

  • All Tools
  • Resume Builder
  • Career Test
  • Pricing
  • For employers

Browse jobs

  • Job Search
  • Jobs by role
  • Companies hiring
  • Remote jobs (US)
  • Jobs in the US
  • Jobs in the UK

Legal

  • Privacy Policy
  • Terms of Service
  • Fair Use Policy

Company

MIRRAI CHAT LTD (Company No. 16403306)

71-75 Shelton Street, Covent Garden

London, WC2H 9JQ, UNITED KINGDOM

contact@mirrai.chat

© 2026 Mirrai Careers. All rights reserved.