Principal Systems Engineer
Oracle
We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.
Responsibilities
Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.
Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.
Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.
Serve as the senior escalation point for complex GPU host and repair issues
Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
Participate in on-call rotations and provide support for critical infrastructure issues.
Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
Build and improve AI agents, ensuring safe rollout, execution and monitoring.
Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.
Required Qualifications:
8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.
Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.
Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.
Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.
Strong problem-solving and troubleshooting skills.
Excellent communication and teamwork skills.
Experience with observability tooling, including metrics, logging, dashboards, and alerting.
Experience with AI agents and tooling
Experience leading on-call operations and incident response.
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Preferred Skills:
Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.
Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.
Qualifications
Minimum Job QualificationsEducation and/or Experience:
11 years of experience in computer network administration, database management, systems architecture, systems administration, or related field OR Bachelor's Degree in information technology, computer science, engineering or related field AND 7 years of experience in computer network administration, database management, systems architecture, systems administration, or related field OR Master’s Degree in information technology, computer science, engineering or related field AND 5 years of experience in computer network administration, database management, systems architecture, systems administration, or related field. Job Skills:
Same skills as prior level plus;
Cybersecurity Trends Demonstrated ability in or knowledge of cybersecurity trends, including staying current with industry threats, best practices, and emerging technologies.
Incident Management and Response Demonstrated ability in or knowledge of incident management and response, including timely handling and escalation of incidents to minimize business impact.
Training and Development Demonstrated ability to design and deliver effective training programs to build team and individual capabilities.
Technical Account Management Demonstrated ability to manage technical relationships and ensure successful outcomes with key accounts.
Cloud Architecture Demonstrated ability in or knowledge of cloud architecture, including designing scalable, reliable, and performant cloud services. Preferred Job Qualifications
Education and/or Experience:
12 years of experience in computer network administration, database management, systems architecture, systems administration, or related field OR Bachelor's Degree in information technology, computer science, engineering or related field AND 8 years of experience in computer network administration, database management, systems architecture, systems administration, or related field OR Master’s Degree in information technology, computer science, engineering or related field AND 6 years of experience in computer network administration, database management, systems architecture, systems administration, or related field. Job Skills:
Same skills as prior level
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on us.fitly.work or by calling View phone number on us.fitly.work in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
$79.9k - $187k
We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational...SuggestedTemporary workFlexible hours$79.9k - $187k
...Job Description We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation...SuggestedTemporary workFlexible hours$142k - $194k
...~ Strong understanding of distributed systems and peer-to-peer, node-to-node, and service... ...~ Bachelors degree in Computer Science, Engineering, or a related field, or equivalent... ...collaborating with hackajob to recruit a Principal Systems Engineer for Oracle Cloud Infrastructures...SuggestedFull time$114.6k - $234.6k
...infrastructure delivery continues to grow, OCI needs hands-on systems thinkers who can operate close to the field, understand where deployment... ...friction into durable improvements. The Field Systems Engineer is a senior individual contributor role focused on improving deployment...SuggestedTemporary workFlexible hoursShift work$114.6k - $234.6k
...per year Requirements: We are seeking 7+ years of experience in operational excellence, field engineering, deployment execution, technical program management, systems or process transformation, business operations, or a similarly complex execution setting. We need...SuggestedFull timeFlexible hoursShift work- ...JOB SUMMARY: The System Engineering Manager provides leadership and technical oversight for BrightRidge's System Engineering Department, ensuring the safe, reliable, and efficient operation of the electric transmission and distribution system. This position manages...Work at officeLocal area
- ...and commercializing energy technologies. If you are searching for the best new ideas and share our vision, join us as a Systems Engineering Manager . This is what you need to know: Location: Knoxville, TN Salary: Highly Competitive Plus Benefits...Permanent employmentFull timeContract work
$114.6k - $234.6k
...Job Description The ideal candidate has strong power-systems engineering experience, particularly in utility interconnection, substations, energy storage, on-site generation, and large-load infrastructure. Key Responsibilities Power Systems Engineering Develop...Temporary workFlexible hours- ...We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational...Flexible hours
$84.9k - $209.5k
This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure...Temporary workFlexible hours$96.3k - $264.1k
...and Technical LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms... ...multiple teams. Serve as a senior technical authority for system reliability, scalability, resilience, performance, and...Temporary workFlexible hours$169.3k - $304.7k
...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,... ..., or DevOps role working with large-scale distributed systems Have a deep understanding of TCP/IP, BGP, load balancing...Work experience placementWork at office- This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle’s relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require...Full timeRelocationRelocation package
$114.6k - $234.6k
...large-scale, highly available distributed systems, with Linux development and systems... ...Compute, Networking, Security, Data Center Engineering, and Hardware Development teams to launch... ...professionals with Oracle. This Principal Member of Technical Staff role is part of...Full timeFlexible hours$114.6k - $234.6k
*This position is based onsite in Nashville, TNWe are seeking a Principal Core Infrastructure Engineer to help design, build, and operate the foundational systems behind a highly available, durable, and globally distributed object storage service. In this senior individual...Temporary workImmediate startFlexible hours- ...hired anywhere in the continental U.S. Optiv is seeking a Principal SailPoint Engineer to join Optiv Security’s 24x7x365 Security Operations Center... ...procedures, as well as managing and maintaining security systems across internal and client environments. The Principal...Full timeWork at officeLocal areaRemote workWork from home
$114.6k - $234.6k
...the job remains posted.Career Level - IC4Qualifications 10+ years of experience in power system modeling, analysis, and studies. Master’s degree or higher in Electrical Engineering, preferably with a power systems focus. 5+ years of recent experience developing PSCAD models...Temporary workFlexible hours- ...us and create a higher standard for a better world. The Principal Engineer provides general technical and engineering management expertise... ...application software; Knowledge of software and database systems. WORKING CONDITIONS AND PHYSICAL REQUIREMENTS Standard business...Local areaRemote work
- ...Overview Our Federal Services Group is seeking a Principal Engineer to join the team. ENERCON Federal Services provides engineering design... ...engineering expertise to the design and qualification of structures, systems, and components Producing procurement specifications and...Contract workFor contractorsRemote work
$121.5k - $264.1k
...service and shares guidance on practices for reliability and functionality. Provides direction to ensure accurate forecasting and ensure systems have adequate resources, identifying resource gaps. Maintains a collaborative relationship with the software development team to...Temporary workImmediate startFlexible hours- ...Principal Engineer Anywhere Type: Contract-to-Hire Category: Development Industry: Technology Workplace Type: Remote... ...Demonstrated ability to influence technical direction across teams and systems. ~ Experience partnering with Product and business...Hourly payContract workLocal areaRemote work
- ...Senior Vacuum Systems Engineer Join an advanced technology programme developing next-generation laser-based uranium enrichment systems. This is a rare opportunity to work at the forefront of nuclear innovation, designing highly specialised vacuum and gas handling systems...Full timeRelocation package
- ...benefits, professional development, and a passionate team of co-workers! Novatech has an exciting opportunity for you as a SYSTEMS ENGINEER , supporting our greater Nashville market. Our Systems Engineer team takes escalations from the Systems Specialist team...Full time
$110k - $165k
...future of our communities. This is a Lead Cloud & Infrastructure Engineering position at the Vice President level, which is part of the job... ...infrastructure and ensuring the seamless operation of IT systems to support business needs effectively.Morgan Stanley is an industry...Temporary workLocal areaRemote workFlexible hours- Functional Systems Engineer Hybrid working - Bristol basedThe OpportunityWe’re working with a leading defence organisation that is looking for an experienced V&V Systems Engineer to join their growing Systems Engineering team.The role will support the continued development...
- ...of co-workers! Novatech has an exciting opportunity for you as a SYSTEMSENGINEER, supporting our greater Nashville market.Our Systems Engineer team takes escalations from the Systems Specialist team and interface with vendors, primary contacts and the clients' IT Staff...Full time
- Systems Engineer Location - Bristol Salary - Up to 57,000 (plus benefits) Copello are seeking a Systems Engineer - Requirements to join the Functional Development team supporting the growing portfolio for a defence organisation based in Bristol on a permanent basis.As...Permanent employment
- Systems EngineerLocation - BristolSalary - Up to 57,000 (plus benefits)Copello are seeking a Systems Engineer to join the Functional Development team supporting the growing portfolio for a defence organisation based in Bristol on a permanent basis. Reporting to the Chief...Permanent employment
- Requisition Id 16927 Overview: We are hiring a HPC Systems Engineer to design, operate and maintain clusters, servers, and workstations supporting services where science happens at ORNL! This position resides in the Emerging Technologies & Computing team in the Research...Work at officeRelocation packageFlexible hours
$68.3k - $148.3k
...performance trend analyses and manages server capacity. Performs system configurations and backups independently. Implements monthly,... ...vendors and cross-functional teams (e.g., Development, Cloud Engineering, Product Engineering, other IT teams) to drive collaboration...Temporary workFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Systems Engineer. Be the first to apply!




