Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Software Engineer (Data Center Automation)

Full-time

x.ai

ABOUT xAI


xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:


We are seeking a highly skilled Sr. Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment. This role focuses on automating processes, building and implementing robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The ideal candidate will combine strong coding abilities with hands-on data center experience to build scalable reliability services, optimize system performance, and minimize downtime—including close partnership with facility operations to address physical infrastructure impacts. If you thrive in lightning-fast, distributed environments and are passionate about leveraging automation to drive efficiency, this is an opportunity to make a significant impact on our infrastructure's resilience and scalability.

In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities. By prioritizing automation and observability, team members in this role can reduce mean time to recovery (MTTR) by up to 50% through proactive monitoring and automated remediation, based on industry benchmarks from high-scale environments like those at hyperscale cloud providers.

The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers. This is achieved through proactive automation, robust observability, and integrated software-physical reliability strategies, ensuring our AI infrastructure remains resilient, scalable, and at the cutting edge of innovation.

RESPONSIBILITIES:



  • Design, develop, and deploy scalable code and services (primarily in Python and Rust, with flexibility for emerging languages) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning. We value adaptability to new tools and paradigms in the fast-evolving AI space.

  • Implement and maintain observability tools and practices, such as metrics collection, logging, tracing, and dashboards, to provide real-time insights into system health across multiple data centers—open to innovative stacks beyond traditional ones like ELK.

  • Collaborate with cross-functional teams—including software development, network engineering, site operations, and facility operations (critical facilities, mechanical/electrical teams, and data center infrastructure management)—to identify reliability bottlenecks, automate solutions for fault tolerance, disaster recovery, capacity planning, and physical/environmental risk mitigation (e.g., power redundancy, cooling efficiency, and environmental monitoring integration).This role encourages broad skill sets from diverse technical backgrounds to foster innovation.

  • Troubleshoot and resolve complex issues in data center environments, including hardware failures, environmental anomalies, software bugs, and network-related problems, while adhering to reliability principles like error budgets and SLAs.** Key Insight: By applying SWE rigor to troubleshooting, team members can create reusable diagnostic tools that accelerate resolution, turning unscheduled events (e.g., hardware faults) into opportunities for system hardening and reducing overall end-user impact through targeted SLAs that prioritize critical AI services. We seek versatile problem-solvers who adapt to bleeding-edge challenges.

  • Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration (e.g., Kubernetes or emerging alternatives), and scripting for automation.

  • Understand network topologies and concepts in large-scale, multi-data center environments to effectively troubleshoot connectivity, routing, redundancy, and performance issues; integrate observability into data center interconnects and facility-level controls for rapid diagnosis and automation.** Key Insight: In multi-site setups, network insights allow for automated failover mechanisms that handle both digital and physical disruptions, ensuring seamless continuity for end-users during events like fiber cuts or power outages. This attracts candidates from varied networking and systems backgrounds to drive forward-thinking solutions.

  • Participate in on-call rotations, post-incident reviews (blameless postmortems), and continuous improvement initiatives to enhance overall site reliability, including joint exercises with facility teams for physical failover and recovery scenarios. We prioritize growth-minded individuals who embrace evolving practices.

  • Mentor junior team members and document processes to foster a culture of automation, knowledge sharing, and adaptability to new technologies.

BASIC QUALIFICATIONS:



  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience).

  • 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering, preferably supporting large-scale, distributed, or production environments.

  • Strong programming skills with proven production experience in Python (required for automation and tooling); experience with Rust or willingness to work in Rust is a plus, but strong coding fundamentals in at least one systems-level language (e.g., Python, Go, C++) are essential.

  • Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments.

  • Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes (or similar systems).

  • Experience implementing observability solutions, including metrics, logging, tracing, monitoring tools (e.g., Prometheus, Grafana, or alternatives), alerting, and dashboards.

  • Familiarity with troubleshooting complex issues in distributed systems, including software bugs, hardware failures, network problems, and environmental factors.

  • Understanding of networking fundamentals (TCP/IP, routing, redundancy, DNS) in large-scale or multi-site environments.

  • Experience participating in on-call rotations, incident response, post-incident reviews (blameless postmortems), and reliability practices such as error budgets or SLAs.

  • Ability to collaborate effectively with cross-functional teams (software engineers, network teams, site/facility operations, mechanical/electrical teams).

PREFERRED SKILLS AND EXPERIENCE:



  • 5+ years of experience in SRE or infrastructure roles, ideally in hyperscale, cloud, or AI / ML training infrastructure environments with multi-data center setups.

  • Hands-on experience operating or scaling Kubernetes clusters (or equivalent orchestration) at large scale, including automation for provisioning, lifecycle management, and high-availability.

  • Proficiency in Rust for systems programming and performance-critical components.

  • Direct experience integrating software reliability tools with physical data center infrastructure (e.g., power, cooling, environmental monitoring, facility controls) and automating responses to physical events.

  • Exposure to advanced or innovative observability stacks beyond traditional tools (e.g., exploring cutting-edge alternatives for metrics, logs, and tracing).

  • Experience building automated remediation, fault tolerance, disaster recovery, capacity planning, or predictive failure detection systems.

  • Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments.

  • Demonstrated success reducing downtime, MTTR, or improving resource efficiency (e.g., through automation or observability) in high-stakes production settings.

  • Prior work with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms.

  • Mentoring experience, strong documentation skills, and a track record of fostering knowledge sharing and automation culture.

  • Comfort with rapid technology adaptation in fast-evolving domains like AI infrastructure.

xAI is an equal opportunity employer. For details on data processing, view our 

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Sr. Software Engineer (Data Center Automation) in Remote vacancy
  • $165k - $242k

     ...platforms. We're standing up new data centers at an extraordinary pace,...  ...platform that gives network engineers, fleet engineers, and...  ...The goal is to build bespoke software to model our infrastructure...  ...drive planning, coordination, automation, of some of the most advanced... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    1 day ago
  • $165k - $180k

     ...and ensure the integrity of networks, data, systems, and processes. Organizations...  ...risks in real time. Join ExtraHop as a Sr. Software Engineer and help us build a fault-resilient,...  ...ExtraHop empowers Security Operations Centers (SOCs) to detect, investigate, and remediate... 
    Senior
    Full time
    Remote work
    Flexible hours

    ExtraHop

    Remote
    a month ago
  • $160k - $175k

     ...the integrity of networks, data, systems, and processes. Organizations...  ...Position SummaryAs a Senior Software Engineer on the Framework team, you...  ...detection, visibility, and automation at scale. This role is ideal...  ...Security Operations Centers (SOCs) to detect, investigate... 
    Senior
    Remote work

    ExtraHop Networks

    Seattle, WA
    1 day ago
  • $170k - $220k

     ...Job Title: Senior Software Engineer – Cloud Infrastructure / Network Automation Industry: Cloud Infrastructure, High-Performance Computing (HPC), Software...  ...Infrastructure-as-Code tools such as Terraform or Ansible Data center networking concepts TCP/IP, DNS, load... 
    Senior
    Full time
    Remote work
    Relocation package

    Addison Group

    Dallas, TX
    8 days ago
  • $110k - $130k

     ...enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of...  ...Vultr is seeking a highly skilled and experienced Senior Software Engineer, Internal Platforms (Data & Domain) to serve as the domain... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Vultr

    United States
    1 day ago
  •  ...building industry-leading martech and data products for the PropTech space. Our focus...  ...MavenAI, delivers AI-powered marketing automation built by industry experts and...  ...The Role We are looking for a Senior Software Engineer to drive the evolution of our shared data... 
    Senior
    Temporary work
    Remote work
    Work from home
    Work visa
    Flexible hours
    Shift work

    ApartmentIQ

    Madison, WI
    a month ago
  • $192k - $240k

     ...management, bill pay, and travel software, Brex enables founders and...  .... Brex’s AI-native automation and world-class service eliminate...  ...you need to grow your career.Engineering at BrexEngineering at Brex is...  ...intention. Our teams span Software, Data, Security, and IT, and... 
    Senior
    Work at office
    Remote work
    Work from home

    Brex

    Seattle, WA
    1 day ago
  • $117.57k - $165.62k

     ...combines analog, digital, AI, and software technologies into solutions...  ...help drive advancements in automation and robotics, mobility, healthcare, energy and data centers. With revenue of more than $11...  ...: Senior Software Development Engineer (Embedded Software)Job Requisition... 
    Senior
    Permanent employment
    Full time
    Work at office
    Remote work
    Work from home
    Day shift
    2 days per week

    Analog Devices

    Wilmington, MA
    1 day ago
  • $104k - $156k

     ...OverviewIn this critical role, the Sr. Field Application Engineer , will work closely with customers,...  ...strategies for applications in Industrial Automation and Electrification industries....  ...Connectivity’s valued customers in the data center infrastructure for critical power... 
    Senior
    Local area
    Remote work

    TE connectivity

    Columbus, OH
    4 days ago
  •  ...machine learning, and data-intensive applications...  ...Mirantis empowers platform engineering teams to deliver...  ...or in sovereign data centers. As enterprises navigate...  ...Mirantis delivers the automation, GPU orchestration, and...  ...Design and build the software that provisions, integrates... 
    Senior
    Full time
    Local area
    Remote work

    Mirantis

    Remote
    22 days ago
  •  ...thinking experts. We're one of the largest engineering and system integration firms in the...  ...providing value for our clients through IT automation and control solutions for more than 30...  ...Experience with PLC or DDC controlled data center projects. Experience with AVEVA... 
    Senior
    Local area
    Remote work
    Flexible hours

    E Tech Group

    Eagle Mountain, UT
    more than 2 months ago
  • $113.01k - $202.16k

    DescriptionWe are seeking a highly experienced Senior Data Center & Cloud Networking Engineer to join our team in a hybrid capacity (3 days onsite /...  ...including hybrid cloud and data center fabricsDrive network automation and operational efficiency initiatives where... 
    Senior
    Traineeship
    Local area
    Remote work

    Mount Sinai Health System

    New York, NY
    1 day ago
  •  ...applications and tools for the effective management and operation of big data platform  • Analyze and improve stability, efficiency, and...  ...of the platform  • Work closely with Architects, lead engineers and business on product design and features.  • Mentor and coach... 
    Senior
    Full time
    Work experience placement

    Jobsbridge

    Remote
    1 day ago
  • $155k - $212k

     ...revolutionizing healthcare by leveraging data and automation to empower care providers (building on...  ...offers a unique opportunity to join a software development team to advance our...  ...As a core member of our software engineering team, you will design, build, and scale... 
    Permanent employment
    Full time
    Home office
    Flexible hours

    Xealth

    Remote
    1 day ago
  • $286.2k - $326.7k

    Sr. Distinguished, Software Engineer - Enterprise Data Storage and Consumption Platforms - Remote-Eligible Job Description As a Sr. Distinguished...  ...or GCP), containerization (Docker, Kubernetes) and automated deployment Capital One is open to... 
    Senior
    Full time
    Part time
    H1b
    Local area
    Remote work

    Capital One Financial Corporation

    Remote
    1 day ago
  • $160k - $225k

     ...the ultimate goal of enabling human life on Mars. SR. SOFTWARE ENGINEER, PROPULSION SIMULATION & DATA ANALYSIS (RAPTOR) The Raptor Systems Modeling and...  ...improvements to manufacturing and engine hardware design Automate processes and software tools that are used to... 
    Senior
    Permanent employment
    Full time
    Temporary work

    Spacex

    Remote
    1 day ago
  • $193.93k - $291.15k

     ...high-output generalists where ML and systems engineering converge to push autonomy performance forward. As a Senior Perception ML Data Infrastructure Engineer, you will own the...  ...You ~ BS/MS in Computer Science , Software Engineering, Robotics, or a related... 
    Senior
    Full time

    Nuro

    Remote
    1 day ago
  •  ...the world's most demanding AI clouds. Senior Software Engineer About Netris Netris is the leading network automation platform for AI clouds. AI is the largest infrastructure...  ...backed by Andreessen Horowitz. SDN reinvented data center networking, and Netris is doing it for the AI... 
    Senior
    Flexible hours

    Netris

    Santa Clara, CA
    2 days ago
  • $109k - $160k

     ...About The Role: The Data Platforms Team serves...  ...innovative solutions, automation and operations of our...  ...are seeking a senior engineer with specialization in...  ...years of experience in a software or infrastructure...  ...in our office and data center locations ~ A casual... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Washington DC
    1 day ago
  • $140k - $190k

     ...Qumulo's cloud data platform manages exabytes of the world's most...  .... About The Position Engineers working on core services build...  ...and protect their data. The software stack includes a distributed...  ...and object storage across data centers, edge, and public clouds,... 
    Full time
    Local area
    Flexible hours

    Qumulo

    Remote
    1 day ago
  • $180k - $210k

     ...reliability and performance. Our data centers are optimized for AI...  ...Role: Crusoe Cloud Network Engineering team is looking for an ambitious...  ...team player. As a Software Engineer, you will be part of...  ...Conceptualize, build, and maintain automation and tools to support New... 
    Senior
    Full time
    Temporary work
    Work experience placement

    Crusoe

    Remote
    1 day ago
  • $98.5k - $141.5k

     ...futures. THE ROLEIn conjunction with the Investment Data Management Office, the Senior Data Engineer contributes to a long-term strategic initiative to unify...  ...lead to design, develop, implement and deploy new software components to investment data platformPartner with data... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    MFS Investment Management

    Boston, MA
    5 days ago
  • $142.2k - $213.4k

     ...Systems sector has an opening for a Sr. Principal Airworthiness Engineer to join our team of qualified, diverse...  ...to come.In an increasingly data-driven world, organizations must develop...  ...management of data within the AWWSC Center of Technical Excellence (CoTE), to assure... 
    Senior
    Full time
    Immediate start
    Remote work
    Relocation package
    Shift work

    Northrop Grumman

    Melbourne, FL
    5 days ago
  • $125k - $150k

     ...SOFTWARE ENGINEER - DATAPATH  Turn a customer’s intent into a live, multi-cloud network — automatically...  ...When a company connects its data centers, clouds, and offices on Alkira , our...  ...Orchestration team, and we build that automation. Using a mix of Ansible, Python, and... 
    Full time
    H1b

    Alkira

    Remote
    1 day ago
  •  ...OVERVIEWThe Senior Information Security Engineer - OT/IoT Security is responsible for...  ...critical infrastructure systems supporting data center operations, including building...  ...experience (Python, PowerShell, etc.) for automation and data analysis.Relevant certifications... 
    Senior
    Remote work

    Quality Technology Services

    Suwanee, GA
    5 days ago
  •  ...forward-thinking infrastructure team, the full-time Senior Data Center Infrastructure Engineer will build, deploy, and optimize on-premises and cloud-...  ...modernization projects while ensuring performance, automation, and reliability. Key responsibilities Build and sustain... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    5 days ago
  • $180k - $250k

     ...Software Engineer, Data Engineering United States (remote) The race to AI has become the race to power. Every breakthrough in artificial...  ...interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving... 
    Remote work
    Flexible hours

    Verse

    United States
    6 hours ago
  • $110k - $185k

     ...We require 4+ recent years of data migration experience within...  ...experience with PLM, ERP, MES, or MRO software. DoD experience is...  ...in Computer Science, Computer Engineering, Information Technology,...  ...worldwide, and our people are at the center of everything we do. We... 
    Senior
    Full time
    Remote work
    Worldwide

    eQ Technologic Inc.

    Seattle, WA
    3 days ago
  •  ...and Military expertise to every case. We are hiring a Software Engineer (Data Management) in Linthicum Heights, MD. Position location is...  ...and authorization process to a new model that emphasizes automation, streamlined processes and approvals, continuous monitoring... 
    Local area
    Work from home
    Flexible hours
    Shift work

    Themis Insight LLC

    Linthicum Heights, MD
    5 days ago
  •  ...Job Description Job Description Job Summary The Sr. Software Engineer - External Data Integrations will code, test, and document software solutions...  ...engineering practices, including source control, automated testing, code review, CI/CD, and secure development principles... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Fidelity & Guaranty Life Insurance Company

    Des Moines, IA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Software Engineer (Data Center Automation). Be the first to apply!