LandMe

Software Engineer, Production Inference (Distributed Inference)

Thinking Machines Lab · San Francisco, California

$350,000–$500,000/yrOnsiteFull-timeDirect from the employer

Posted Oct 9, 2026 · Verified open Oct 9, 2026

Apply on the employer's siteFind more jobs like this on LandMe

About the job

ABOUT THINKING MACHINES

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

ABOUT THE ROLE

We're hiring a Software Engineer to build and scale the distributed production inference systems that serve Inkling, Inkling-Small, and Tinker in production. You'll own the systems that turn trained models into fast, reliable, cost-efficient services — from request routing and batching to multi-node serving and GPU utilization at scale.

This is a systems-heavy, production-first role. You'll work closely with research and infrastructure teams to translate rapidly evolving model architectures into serving systems that meet real-world latency, throughput, and reliability requirements, and you'll be on the front line when production inference systems need to scale, recover, or improve.

WHAT YOU'LL DO

SKILLS & QUALIFICATIONS

PREFERRED QUALIFICATIONS

LOGISTICS