Senior Site Reliability Developer
Oracle
:
We are seeking associate passionate about automation, cloud computing and application security. Work with application delivery teams on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Work closely with agile teams to ensure they have the tools needed to write, test and deploy code with ease and ensure dev and QA satisfaction and needed Security reviews.
Plays a critical role in providing observability & Monitoring tooling, Change Automation tooling, security tool integration, automation support, and business continuity program for day-to-day operations.
Qualifications
- BS or MS in Computer Science or closely related field.
- 6+ years of work experience, with 4+ years in DevOps and Observability.
- Certification in any cloud hosting platform such as Google/Azure/AWS/Oracle
Skills:
- Hands-on Experience on Monitoring and Observability tooling development and Configuration.
- Experienced in working with tools like Prometheus, Grafana, AppDynamics, Loki, Alertmanager etc.
- Strong knowledge of Enterprise application architectures hosted on the cloud.
- Knowledge of Infrastructure Automation Tools - Terraform, Chef, Ansible, Saltstack, Puppet, etc (Infrastructure As Code)
- Knowledge of phases of the Software development life cycle
- Knowledge on SW delivery pipeline, development environment, Build & Integration.
- Comprehensive Knowledge of Continuous Integration Skills - Version control, builds, and remediation.
- Strong automation skills (tool agnostic) and the ability to drive initiatives to automate processes
- SW Testing - Unit, Integration, System, Load testing, performance, security, regression testing
- Scripting & Language Skills - Python, Go, Groovy(/Java), Shell, Perl etc
- Software Security Skills: Knowledge of PCI-DSS, HIPPA, SOX, GDPR, and CCPA Standards and Policies and the associated certification and audit processes
- Knowledge of Integration of Software security tools to the CICD process.
- Experience working with applications developed in Java, Node.js, Python, restful web services, and APIs.
- Strong operational experience in Linux/Unix environment
- Experience with REST-style web services / APIs.
- Context-switch between multiple projects/codebases/concepts with ease
- Understand software development at a fundamental level, use the best tools for the job, and always think about the future (at scale) when architecting solutions
- Knowledge of fundamental aspects of release automation (packaging, dependencies, promotion, deployment, compliance)
- Knowledge of the desired tools for release automation, Jenkins, repository management (SVN, GIT), and deploying software through scripts (ANT, Make, Shell script, Golang scripts)
- Experience with technologies like Kafka, Docker, Elasticsearch, continuous integration (Drone, Jenkins, Travis, Bamboo) and understanding its benefits, workflows, etc.
- Experience on project and ticket management tools such as JIRA and insight on quality analysis as well
- Exposure in integrating testing tools such as Selenium, QTest Manager, etc
- Proficient in some of these: Chef, GitHub, DevOps, Dockers, Jenkins, Black chair
- Cloud experience (SaaS and PaaS) on Public Cloud e.g., AWS, Google, Azure, Oracle
- Develop & implement DevOps Solutions in the areas of Observability, Change Automation, and Service Management.
- Handles cross-functional collaboration to develop tools for secure, scalable, and reliable systems.
- Identify, integrate, monitor, and improve infosec controls by understanding business processes.
- Implements DevOps design principles and security best practices using relevant skills and experience.
- Lead & influence the architectural design of features by determining quality & adhering to specifications.
- Design & build solutions that move data from internal solutions to cloud-based solutions.
- Responsible for Architecture/Design comprising, Security, Risk & Compliance
- Monitor the delivery of solutions between architecture, time, cost, and quality and provide an Assessment of Costs and Benefits, and anchor continuous improvement.
- Adopt the OCI standard tools and DevOps processes.
- Tenets and best practices of Continuous Monitoring and Observability solutions.
- Provides solutions for Continuous Delivery and Deployment (CD) and Change Management Automation
- Provides solutions for Continuous Monitoring (CM) - monitoring and analysis of infrastructure, process, and applications
- Provides solutions for Tooling for Checkout, Compile, Package, Verify, Quality Scan, Publish Artifact, Deploy to Development and Automation Testing
- Provides solutions for Build & Release tools on Linux and Windows VMs/ Compute instances and containers.
- Provides Code Quality Metrics, assists the team in improvement strategies
- Performance Testing and tools Release processes
- Ability to gel well with Agile teams in real spirit (culture & mindset)
- Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation, and refinement.
- Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning, and launch reviews.
- Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
- Participate in incident handling and other related duties to support the information security function.
- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
- Practice sustainable incident response and blameless postmortems.
- Collaborate with Agile teams in defining technical requirements and best practices with containerized and cloud-native applications.
- ...with software, platform, and networking teams to improve service reliability and deployment workflowsDeploy and maintain network monitoring,... ...in the on-call rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar roleHave...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...provisioning, upgrades, patching, and deletion.Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and...SeniorWork at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization... ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in...SeniorFull time$174k - $252k
...activities such as system design consulting, developing software platforms and frameworks,... ...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Senior$267k - $356k
...most demanding compute workloads in the industry, which means reliability and performance aren't just goals—they're the baseline. We're looking... ...of software-defined storage across new and existing sites using tools such as Ansible, Jenkins etc.Work with hardware and...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$148k - $235.75k
...Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself (not the...SeniorFull time$101k - $161k
...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... .../Pulumi, Bash. You will be expected to develop, operate, and work with many different... ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$152k - $241.5k
...NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-... ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data...SeniorFull time$98.9k - $228.7k
What you can expectWe are hiring a Senior DevOps Engineer to ensure reliability, scalability, and operational excellence for our real-time communications platform. This platform supports audio/video conferencing, recording, and live-streaming functionalities. The position...SeniorFull timeWork at officeRemote workFlexible hours$192.4k - $275.8k
...of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...— this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the...SeniorFull timeTemporary workLocal areaFlexible hours- ...Job Title: Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must... ...meet FedRAMP High and IL5 requirements without sacrificing developer velocity Support ATO processes, including documentation...SeniorFull time
- ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 8+ years of experience in SRE, DevOps, or platform engineering Hands-on experience...SeniorFull time
$152k - $241.5k
...Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers... ...Chips Simulation as a trusted and reliable virtual platform.What you will be doing:... ...simulation environment interactionsImprove developer experience for teams using simulation infrastructure...SeniorFull time$307k - $427k
...to define the long-term roadmap for spatial intelligence.Mentor senior technical leads and staff engineers, fostering a culture of... ...software, and hardware. Teams across this area research, design, and develop new technologies to make our user's interaction with computing...Senior$152k - $241.5k
...our GPU Software team. In this role, you will help design and develop key components of our production GPU kernel drivers and embedded... ...environmentsCollaborate with globally distributed teams to deliver scalable, reliable, and high-impact GPU software solutionsWhat we need to see: BS...SeniorFull time$184k - $287.5k
Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation...SeniorFull time$152k - $241.5k
The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to the market! This role involves developing and productizing innovative solutions that will transform transportation and the field of self...SeniorFull time$224k - $356.5k
...motivated engineers who bring GeForce NOW to life! As a member of the GeForce Now Platform Engineering team, you will help design, develop, and optimize high performant GeForce Now gaming servers with state-of-the-art NVIDIA GPUs. Are you passionate about working in a cross...SeniorFull time$224k - $356.5k
...environment with fast pace and agility.What you'll be doing:Developing and triaging platform drivers which goes into SOCsBuilding sophisticated... ...debugging is invaluableExperience working on system level reliability and resiliency features.Familiarity with system level...SeniorFull time$152k - $241.5k
Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is an engineering... ...facing GPU cloud services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through...SeniorFull time$184k - $287.5k
...based infrastructure that allow agents to interact with internal developer systems (Gerrit, NVBugs, CI/CD, Perforce, Slack)Design and... ...across GPU SW teams by shipping agentic services that are fast, reliable, and meaningfully better than the manual alternative!What we...SeniorFull time$184k - $287.5k
We are now looking for a Senior System Software Engineer to work in our Tegra system software group. The best candidates will have excellent... ...If you're a creative software engineer with a real passion for developing products with new technology, we want to hear from you. Join us...SeniorFull time$152k - $241.5k
...Proton titles.Identifying changes to API usage to improve performance and communicating via appropriate channels with third-party developers.Implementing driver performance improvements and resolving driver defects.Collaborating with engineers on the team and across...SeniorFull time$184k - $287.5k
...lower latency in batch ETL workloads. Apache Spark is the most popular data processing engine in data centers. At NVIDIA, we are developing an open source plugin to accelerate Spark applications on GPUs without any code changes.What you'll be doing:Enable C++ native execution...SeniorFull time$184k - $287.5k
...supercomputing, gaming, and visualization. As a Senior System Software Engineer on the NvSci... ...and maintainability, and elevate developer experience.Evaluate trade-offs in resource... ...generative AI technologies to improve software reliability, maintainability, and scalability.What...SeniorFull timeRemote work$170k - $277k
...assisted coding tools (e.g., Cursor, Claude); you optimize your workflow and the team’s workflow around them.Strong desire to mentor and develop junior engineering talent, with a demonstrated ability to transition from a high-impact individual contributor to a technical...SeniorFull timeRemote workVisa sponsorshipWork visa$184k - $287.5k
...product lines and OEMs. We are looking for a highly motivated Senior Software Systems Engineer with a strong foundation in software... ...hard-working attitude.Ways to Stand Out From The Crowd:Experience developing ADAS software.Deep understanding of real-time operating systems...SeniorFull time$184k - $287.5k
...and efficiency for our next generation of datacenter products, including CPUs and CPU+GPU Superchips. What you will be doing:Design, develop, test, and optimize software for our next-generation SoCs. In both pre-silicon and post-silicon phases of execution.Review...SeniorFull timeRemote work$152k - $241.5k
NVIDIA Solutions Engineering team is searching for engineers to help develop and bring NVIDIA’s autonomous vehicle platform to the world. You will work on state of the art technologies alongside experts in Deep Learning, Computer Vision, and vehicle control for NVIDIA’s...SeniorFull time$152k - $241.5k
...looking for energetic, enthusiastic and technologically savvy engineers to join the Tegra Core Firmware team. Here we architect and develop the boot stack firmware for the flagship Tegra chipset which are the core components for high-compute platforms for automotive,...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Developer. Be the first to apply!
- site reliability engineer sre Santa Clara, CA
- site reliability engineer Santa Clara, CA
- senior operations technician Santa Clara, CA
- senior cloud service delivery manager Santa Clara, CA
- senior it service manager Santa Clara, CA
- senior chief engineer Santa Clara, CA
- sr operations manager Santa Clara, CA
- senior physical design engineer Santa Clara, CA
- senior account director Santa Clara, CA
- senior director clinical development Santa Clara, CA

