jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

infrastructure

Posted 3 days ago

AI Infrastructure Engineer, Sandbox Platform

at scaleai

London, United KingdomOn-site
You are nearing today's limit. Upgrade for unlimited access.

Responsibilities

  • Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments
  • Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads
  • Optimise for cold-start latency, memory footprint, and resource utilisation at scale
  • Drive down error rates through systematic debugging, monitoring, and proactive fixes
  • Lead architecture reviews and own projects end-to-end, from design through deployment, in fast-paced cross-functional settings Ideally you'd have: 4+ years of

Requirements

  • As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both internal and customer-managed environments.
  • experience of the engineers and researchers using this system as they do about the kernel internals underneath it.
  • experience: clean APIs, clear error messages, good docs, and a client library that feels well-crafted.
  • experience building high-performance systems software, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs
  • Deep understanding of Linux internals: process isolation, memory management, cgroups, namespaces, etc. •
  • Experience with containerisation and virtualisation technologies (e.g., Docker, Firecracker, gVisor, QEMU, Kata Containers)
  • Proficiency in a systems programming language such as Go, Rust, or C/C++
  • Comfort working across infrastructure layers, from kernel modules to orchestration frameworks (e.g., Kubernetes)
  • Strong debugging skills and the ability to navigate performance/security tradeoffs in production systems
  • Experience as a founder or early engineer at an infrastructure-focused startup, owning a product end-to-end
  • Familiarity with LLM agents and agent frameworks (e.g., OpenHands, Agent2Agent, MCP) •
  • Experience running secure workloads in multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)
  • Exposure to snapshotting and restore techniques (e.g., CRIU, VM snapshots, overlays)
  • At Scale, our mission is to develop reliable AI systems for the world's most important decisions.
  • Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
  • We are expanding our team to accelerate the development of AI applications.

Benefits

  • We comply with the United States Department of Labor's Pay Transparency provision .

Contact

  • If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at accommodations@scale.com.

Additional details

  • This is a role for someone who cares as much about the
  • You'll combine deep systems expertise (isolation, virtualisation, performance) with an obsession for developer
  • You'll partner closely with internal teams to understand how they use the platform, debug their issues, and shape a roadmap that balances immediate needs with long-term architecture. You will:
  • Partner closely with internal teams using the platform to understand their needs, debug issues, and build tooling that serves their use cases
  • Respond to incidents and production issues with urgency, conducting root cause analysis and implementing preventive fixes
  • Help develop and maintain a product roadmap for sandboxing, balancing immediate needs against long-term architectural investment
  • experience — API design, error propagation, documentation, and the small details that make a library feel well-crafted
  • History of on-call/incident response for production systems
  • PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants. About Us:
  • We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.

Find more real-time jobs on JobLoom.