Principal Systems Engineer (Codex)
$142k - $194kOracle
Salary: $142,000 - 194,000 per year Requirements:
- 8+ years of software operations or infrastructure automation experience, with strong Python and Bash skills.
- Expert Linux administration experience, particularly with Ubuntu and Oracle Linux in large-scale production environments.
- Strong understanding of distributed systems and peer-to-peer, node-to-node, and service-to-service communication.
- Data-center and host lifecycle experience, including provisioning, validation, repairs, hardware replacement, and fleet recovery.
- Strong troubleshooting and problem-solving abilities, with excellent communication and teamwork skills.
- Experience with observability tools for metrics, logging, dashboards, and alerting.
- Experience with AI agents and related tooling.
- Experience leading on-call operations and incident response.
- Bachelors degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Hands-on experience with NVIDIA- or AMD-based GPU infrastructure and familiarity with GPU shapes or instance families is preferred.
- Experience operating or automating GPU, compute, or other large-scale cloud infrastructure is preferred.
- The posting also lists minimum qualifications of 11 years of relevant experience; or a related bachelors degree and 7 years of relevant experience; or a related masters degree and 5 years of relevant experience.
- Additional listed skills include knowledge of cybersecurity trends, incident management and response, training and development, technical account management, and cloud architecture.
- Preferred qualifications also include 12 years of relevant experience; or a related bachelors degree and 8 years of relevant experience; or a related masters degree and 6 years of relevant experience.
- Lead GPU repair operations, diagnosis, and resolution of complex, high-severity failures across hardware, Linux, networking, drivers, firmware, and dependent cloud services.
- Develop and maintain automation and operational tools for GPU infrastructure across multiple regions.
- Partner with software engineering, hardware, and operations teams to maintain high GPU fleet availability.
- Improve monitoring, alerting, and diagnostics for fleet health, performance, capacity, and utilization, using tools such as Grafana.
- Act as the senior escalation point for complex GPU host and repair issues.
- Respond to incidents, conduct root-cause analysis, and address issues affecting GPU capacity, availability, and regional deployments.
- Improve AI2 Ops processes, GPU fleet automation, and readiness for OCI region builds.
- Participate in on-call rotations and support critical infrastructure incidents.
- Document operating procedures, automation workflows, troubleshooting guides, and runbooks.
- Develop and improve AI agents, overseeing their safe rollout, execution, and monitoring.
- Mentor junior engineers, provide senior technical guidance, and promote continuous improvement.
- AI
- AI Agents
- Bash
- Cloud
- Firmware
- Grafana
- Hardware
- Incident Management
- Support
- Linux
- Oracle
- Python
- Ubuntu
- NodeJS
- Network
More:
We are collaborating with hackajob to recruit a Principal Systems Engineer for Oracle Cloud Infrastructures AI Infra Operations team. In this technical operations leadership role, you will support GPU infrastructure and help scale reliable operations as OCIs GPU fleets grow across regions. You will work closely with engineering and operations partners, with a focus on availability, repair efficiency, automation, observability, and data-driven operational improvement.
last updated 41 week of 2026
$79.9k - $187k
We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational...SuggestedTemporary workFlexible hours$114.6k - $234.6k
...infrastructure delivery continues to grow, OCI needs hands-on systems thinkers who can operate close to the field, understand where deployment... ...friction into durable improvements. The Field Systems Engineer is a senior individual contributor role focused on improving deployment...SuggestedTemporary workFlexible hoursShift work$114.6k - $234.6k
...per year Requirements: We are seeking 7+ years of experience in operational excellence, field engineering, deployment execution, technical program management, systems or process transformation, business operations, or a similarly complex execution setting. We need...SuggestedFull timeFlexible hoursShift work$142k - $194k
...Requirements: At least 5 years of experience in software operations, systems administration, or infrastructure automation, with proficiency... ...bachelors degree in information technology, computer science, engineering, or a related field and 4 years of relevant experience; or a...SuggestedFull time$114.6k - $234.6k
...Job Description The ideal candidate has strong power-systems engineering experience, particularly in utility interconnection, substations, energy storage, on-site generation, and large-load infrastructure. Key Responsibilities Power Systems Engineering Develop...SuggestedTemporary workFlexible hours$84.9k - $209.5k
This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure...Temporary workFlexible hours$96.3k - $264.1k
...and Technical LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms... ...multiple teams. Serve as a senior technical authority for system reliability, scalability, resilience, performance, and...Temporary workFlexible hours$169.3k - $304.7k
...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for:... ...Engineering, or DevOps role working with large-scale distributed systems Have a deep understanding of TCP/IP, BGP, load balancing...Work experience placementWork at office- This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle’s relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require...Full timeRelocationRelocation package
$114.6k - $234.6k
...large-scale, highly available distributed systems, with Linux development and systems... ...Compute, Networking, Security, Data Center Engineering, and Hardware Development teams to launch... ...professionals with Oracle. This Principal Member of Technical Staff role is part of...Full timeFlexible hours$114.6k - $234.6k
...the job remains posted.Career Level - IC4Qualifications 10+ years of experience in power system modeling, analysis, and studies. Master’s degree or higher in Electrical Engineering, preferably with a power systems focus. 5+ years of recent experience developing PSCAD models...Temporary workFlexible hours$114.6k - $234.6k
*This position is based onsite in Nashville, TNWe are seeking a Principal Core Infrastructure Engineer to help design, build, and operate the foundational systems behind a highly available, durable, and globally distributed object storage service. In this senior individual...Temporary workImmediate startFlexible hours- ...hired anywhere in the continental U.S. Optiv is seeking a Principal SailPoint Engineer to join Optiv Security’s 24x7x365 Security Operations Center... ...procedures, as well as managing and maintaining security systems across internal and client environments. The Principal...Full timeWork at officeLocal areaRemote workWork from home
- ...us and create a higher standard for a better world. The Principal Engineer provides general technical and engineering management expertise... ...application software; Knowledge of software and database systems. WORKING CONDITIONS AND PHYSICAL REQUIREMENTS Standard business...Local areaRemote work
- ...Principal Engineer Anywhere Type: Contract-to-Hire Category: Development Industry: Technology Workplace Type: Remote... ...Demonstrated ability to influence technical direction across teams and systems. ~ Experience partnering with Product and business...Hourly payContract workLocal areaRemote work
$121.5k - $264.1k
...service and shares guidance on practices for reliability and functionality. Provides direction to ensure accurate forecasting and ensure systems have adequate resources, identifying resource gaps. Maintains a collaborative relationship with the software development team to...Temporary workImmediate startFlexible hours- ...benefits, professional development, and a passionate team of co-workers! Novatech has an exciting opportunity for you as a SYSTEMS ENGINEER , supporting our greater Nashville market. Our Systems Engineer team takes escalations from the Systems Specialist team...Full time
- ...of co-workers! Novatech has an exciting opportunity for you as a SYSTEMSENGINEER, supporting our greater Nashville market.Our Systems Engineer team takes escalations from the Systems Specialist team and interface with vendors, primary contacts and the clients' IT Staff...Full time
$68.3k - $148.3k
...performance trend analyses and manages server capacity. Performs system configurations and backups independently. Implements monthly,... ...vendors and cross-functional teams (e.g., Development, Cloud Engineering, Product Engineering, other IT teams) to drive collaboration...Temporary workFlexible hoursShift work- Position Summary:The Systems Engineer is the hands-on individual responsible for the design, security, and dayto- day operation of the Wicked Problems Lab's research computing environment — spanning AWS cloud, on-premises GPU/AI compute, databases, and systems security...Local area
- Robert Half is seeking a Contract Systems Engineer to join our client's IT infrastructure team. In this role, you will be responsible for the design, implementation, maintenance, and optimization of the organization’s systems and infrastructure. This contract position...Contract work
- ...DescriptionMechanical Subject Matter Expert/Principal EngineerStep into a pivotal role at... ...clients include universities, healthcare systems, state, local, and federal government agencies... ...is accelerating demand for our engineering expertise, we are seeking a Mechanical SME...Contract workFor contractorsWork at officeLocal area
$125k - $150k
Job DescriptionRequirements:5+ years of progressive experience in technical consulting, systems engineering, or technical specialist role, performing solution design and customer-facing responsibilities.3+ years of technical sales experience in a relevant industry, with...Full timeShift work$177.6k - $244.2k
...part of a culture that values trust, accountability, and shared success where your work truly matters.Job SummaryThe District Systems Engineer, Commercial, is a vital part of our sales team, serving as a trusted technical advisor to customers and helping them secure their...Full timeRemote workVisa sponsorshipWork visa$114.6k - $234.6k
The Principal Optical Engineer will be joining the Optical Network Engineering team which is responsible for building and scaling Oracle's Optical... ...well as FOA deployment of optical transport solutions and systems. The successful candidate will have a direct impact on our...Temporary workImmediate startFlexible hours$114.6k - $234.6k
...begins architecting components of scalable, elastic distributed systems for the Network Automation Team. Defines and enforces... ...within change‑management plans.We are seeking a Core Infrastructure Engineer to design, build, and operate the distributed systems that power...Temporary workFlexible hours$114.6k - $234.6k
As a Principal Member of Technical Staff, you will own the software design and development for major components of... ...lead developer, curious problem solver, a distributed systems generalist and/or skilled Linux engineer with Systems triage experiance able to dive deep...Temporary workFlexible hours$114.6k - $234.6k
Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms...Temporary workFlexible hoursShift work$114.6k - $234.6k
As a Principal Core Infrastructure Engineer within the Networking organization, you will spearhead the development of innovative services in the Networking... ...an experienced engineer with a deep understanding of systems design and distributed systems to build scalable, high-...Temporary workLong distanceFlexible hours$140k - $210k
...(*Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage... ...a team that applies software engineering principles to make systems more reliable and scalable. These systems support Indeed’s mission...Work experience placementLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Systems Engineer (Codex). Be the first to apply!
- chief engineer Nashville, TN
- senior chief engineer Nashville, TN
- principal infrastructure engineer Nashville, TN
- principal developer Nashville, TN
- general engineer Nashville, TN
- director software engineering Nashville, TN
- engineering director Nashville, TN
- director data engineering Nashville, TN
- principal cloud engineer Nashville, TN
- hotel chief engineer Nashville, TN




