Staff Infrastructure Software Engineer, Fleet & Automation
Nscale
Staff Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role We're hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale's fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments — driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world's leading neo-cloud providers. This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform. You will work closely with teams across Nscale — including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering — to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it. NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback. What You'll Be Doing Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability. Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure. Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale. Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work. Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act. Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement. Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation. Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions. Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them. Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization. Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety. About You 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems. A Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience. A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging. Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services. Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management. Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority. Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis. Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment. A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution. Strong Candidates Will Have Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes. Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA. Experience with bare-metal lifecycle automation and infrastructure management tools such as Redfish, IPMI, PXE, MAAS, Ironic, NetBox, DCIM, OpenStack, or equivalent systems. Experience with workflow orchestration and reliable automation systems such as Temporal, Airflow, Prefect, or event-driven architectures. Experience with Kubernetes, containers, infrastructure as code (Terraform, Pulumi, Ansible), and public-cloud or private-cloud platforms. Experience with observability platforms and high-cardinality telemetry, such as Prometheus, Grafana, OpenTelemetry, ELK, or equivalent. Experience with hardware qualification, burn-in, validation, fleet health, or automated remediation for servers, GPUs, or network equipment. A track record of technical leadership: setting direction, defining reusable patterns, developing other engineers, and improving the effectiveness of multiple teams. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale
- Introduction At IBM Software, we transform client challenges into... ...operates the foundational infrastructure layer that powers Confluent... .... About the Role As a Staff Software Engineer on the Secure Compute Platform... ...operates across a large fleet of clusters spanning multiple...FleetRemote work
$184k - $276k
...in. Platform Engineering builds what the rest... ...engineering, automation, developer tooling... ...systems that turn a fleet of data centers... ...contributor role. A staff engineer is not a... ...of professional software engineering... ...close to physical infrastructure, where the software...FleetDaily paidTemporary workLive outRelocation package- ...Stripe is a financial infrastructure platform for... ...let every Stripe engineer ship code, configuration... ...the Resource Automation and Feature... ...how Stripe ships software. The deployment platform... ...strategy. As a Staff engineer on... ...migrating large fleets from VM-based to...FleetWork at officeLocal areaEarly shift
$231k - $267.5k
...you are Metropolis is seeking a Senior Staff Software Engineer, Recognition Platform to serve as the... ...Deep understanding of observability, automated testing, progressive delivery, and fluency... ...Experience with IoT, edge computing, fleet-scale device telemetry, or MQTT-class...FleetTemporary workWork at officeLocal area- ...will work alongside a team of strong software engineers and act as a force multiplier for our... ...If you want to learn more about our ML Infrastructure, here is one of our past talks at re:Invent... ...ground-up, fully autonomous vehicle fleet and the supporting ecosystem required...Fleet
$75k - $215k
...Position Description GEICO is seeking an experienced software engineer to work on our Infrastructure Automation Tools team to help drive our tech transformation... ...private cloud, our network, and our hardware fleet. Managing a private cloud in an enterprise environment...FleetHourly payFull timeWork experience placementLocal areaRemote work$201k - $315k
Staff Software Engineer Zoox is looking for an experienced Staff Software Engineer to build, scale... ...our custom High-Performance Computing infrastructure. As Zoox scales its autonomous... ...ground-up, fully autonomous vehicle fleet and the supporting ecosystem required...FleetTemporary workRelocation package$320k
Staff+ Software Engineer, Caching San Francisco, CA | New York City, NY | Seattle, WA About Anthropic... ...fast and correct: a managed Redis fleet, client libraries used across the company... ...the technical direction for caching infrastructure used across Product and Research...FleetWork at officeVisa sponsorshipFlexible hours$320k
Staff + Sr. Software Engineer, Scaling New York City, NY; San Francisco, CA; Seattle, WA About Anthropic... ...from intelligent request routing to fleet-wide orchestration across diverse AI... ...the high-performance inference infrastructure they need to develop next-generation...FleetWork at officeWorldwideVisa sponsorshipFlexible hours$200k - $260k
Staff Software Engineer (Scopely Explore, Inc.; fka Niantic, Inc., San Francisco, California): Research, design, and develop computer and network... ...!” and “Pokémon GO,” along with “Stumble Guys,” “Star Trek™ Fleet Command,” “MARVEL Strike Force,” “WWE Champions,” the...FleetLocal areaImmediate startRemote workWorldwide- ...the future. DataRobot's Fleet team is the engine behind how our platform runs... ...foundational Kubernetes infrastructure that powers everything... ...and intelligent platform automation-while lowering Total Cost... ...where you come in. As a Staff Software Engineer, you'll be responsible...FleetFull timeLocal areaRemote workWorldwideFlexible hours
$182.4k - $247k
...running the world's best data and AI infrastructure platform, so our customers can... ...SaaS companies in the world. Our engineering teams build highly technical products... ...and operate one of the largest scale software platforms. The fleet consists of millions of virtual machines...FleetWork at officeLocal areaWorldwideFlexible hours$140k - $193k
Staff Software Engineer - Backend Hybrid - San Francisco, California; Hybrid... ...vehicles, equipment, and fleet related spend in a single... ...reduces manual workloads by automating and simplifying tasks. Motive... ...work on our AWS Cloud infrastructure, and mentor and learn from...FleetTemporary workWork experience placementWork at office2 days per week$200k - $285k
...Staff Software Engineer - Agent Simulation Tech Lead / Manager Model and simulate human behaviors, specifically traffic and other pedestrian... ...is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology...FleetTemporary workRelocation package$209.1k - $282.9k
AI Compute Infra Engineer As an engineer on the AI Compute Infra... ...build, and operate large-scale infrastructure for AI training, fine-tuning... ...their workloads and automate cluster provisioning, upgrades... ...developing reliable infrastructure software. • Practical knowledge of...Work at officeLocal areaVisa sponsorshipRelocation package- Staff Storage Software Engineer Lambda, the superintelligence cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous...Local areaFlexible hours
- ...Limited in Seattle is searching for a Robotics Software Engineer to design algorithms and systems that enhance the efficiency of our robotic fleet. You will ensure seamless operation and communication within our advanced automation technologies. The ideal candidate has a...Fleet
$201.3k - $302.2k
Staff Software Engineer, ASE Storage Infrastructure Apple Services Engineering (ASE) designs, builds, and operates the cloud infrastructure, server systems,... ...coding, scrubbers, silent-corruption detection, and automated repair/reconstruction. Track record of driving cross...Relocation$217.2k - $288.4k
...development. We do this by building and running the world's best data and AI infrastructure platform, so our customers can focus on the high-value challenges that are central to their missions. Our engineering teams build highly technical products that fulfill real, important...Local areaWorldwide$157.9k - $213.6k
We are seeking a Hardware Engineer III (Mechanical) to lead the mechanical... ...on delivery van and station infrastructure modifications. This role owns... ..., and ease of maintenance at fleet scale.Support pilot... ...employees, supervisors, and staff; adhere to standards of excellence...FleetLocal areaFlexible hours$229k - $343k
.... We are looking for an L6 Staff Tech Lead to join our team responsible... ...the critical Service Infrastructure that powers Snap's entire... ...This role is for a backend engineer focused on Service & Compute... ...practical experience. 9+ years of software development experience; or a...Full timeLive inWork at officeLocal area$79.2k - $209.5k
...doing it with a small group of engineers concentrated on high-... ...wrong action taken across a fleet is worse than no action at all... ...Oracle brings together the data, infrastructure, applications, and expertise... ...professional experience in software developmentDemonstrated...FleetFull timeTemporary workFlexible hours$170k - $250k
...Founded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the... ...TaxBit is looking for a Staff Software Engineer to join our Systems, Security... ...take ownership of the cloud infrastructure that powers our platform. In this...Full timeWork at officeWork from homeFlexible hours- Introduction At IBM Software, we transform client challenges into solutions. Building... ...potential, we are looking for a Staff Software Engineer 2 to join our Traffic team which is responsible... ...not limited to) networking, cloud infrastructure, and multi-tenancy. Independently...Work experience placement
$320k
Staff+ Software Engineer, Infrastructure (Distributed Systems) San Francisco, CA | New York City, NY | Seattle, WA About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and...InternshipWork at officeVisa sponsorshipFlexible hours$209.1k - $282.9k
As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction... ...services, Kubernetes, networking, and compute infrastructure. Turn incidents and operational... ...workload lifecycle management. Strong software and production engineering skills , including...Work at officeLocal areaVisa sponsorshipRelocation package$207k - $300k
...and developing large-scale infrastructure, distributed systems or networks... ...Master’s degree or PhD in Engineering, Computer Science, or a... .... Understanding of hardware/software boundaries, including performance... ...software solutions. As a Staff Software Engineer on the Google...Temporary workWorldwide$198k - $326k
Sr. Staff Software Engineer - Systems Infrastructure This role will be based in Mountain View, CA, or Bellevue, WA. At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The...Work at officeFlexible hours$154k - $220k
...? Join us at Zscaler. Role We are looking for a Sr. Staff Software Development Engineer-AI Security to join our team. This is a Hybrid (based... ...will be responsible for designing and implementing core infrastructure components and distributed systems, serving as a...Full timeWork at officeLocal area- Staff + Sr. Software Engineer, Cloud Inference Launch Engineering San Francisco, CA | Seattle, WA About... ...intact. This is high-leverage infrastructure work: validation has to be fast and... ...users Have a track record of building automation or test infrastructure that...Work at officeVisa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Infrastructure Software Engineer, Fleet & Automation. Be the first to apply!
- lead infrastructure engineer Seattle, WA
- security infrastructure engineer Seattle, WA
- principal infrastructure engineer Seattle, WA
- entry level infrastructure engineer Seattle, WA
- data infrastructure engineer Seattle, WA
- infrastructure engineer Seattle, WA
- remote infrastructure engineer Seattle, WA
- senior infrastructure engineer Seattle, WA
- infrastructure developer Seattle, WA
- infrastructure engineering manager Seattle, WA




