Location: San Francisco or Croatia
Type: Full-time
Reports to: Engineering Lead, GPU Sandboxes
The role
You will build the runtime that gives every AI agent its own GPU. Full passthrough, hard isolation, serverless economics. This is deep systems work: hypervisors, kernels, drivers, and the orchestration layer that makes an GPUs appear behind an API call in seconds.
What You'll Do
- Build and harden the GPU sandbox runtime: KVM/QEMU, VFIO passthrough, IOMMU group management, guest driver and CUDA stack lifecycle
- Attack cold-start latency across the stack: image caching, VM boot paths, driver init, snapshot restore
- Build fleet-level systems: GPU-aware scheduling and bin-packing, health monitoring, failure detection, automated remediation
- Implement pause, snapshot, and fork semantics for GPU-attached workloads
- Own security-relevant surfaces: device isolation, side-channel considerations, tenant boundaries on shared hosts
- Debug the ugly stuff: PCIe errors, IOMMU faults, driver crashes, thermal throttling, the works
What We're Looking For
- 5+ years of systems engineering in production environments
- Hands-on depth in Linux virtualization internals: KVM, QEMU, libvirt, VFIO, IOMMU
- Strong Go, Rust, or C/C++ and real Linux kernel-adjacent debugging skills
- Experience operating GPUs in production (NVIDIA data center hardware preferred)
- You ship, you own your services, and you like hard problems more than big teams