Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Developer

Full-time

Oracle

:

We are seeking associate passionate about automation, cloud computing and application security. Work with application delivery teams on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Work closely with agile teams to ensure they have the tools needed to write, test and deploy code with ease and ensure dev and QA satisfaction and needed Security reviews.

Plays a critical role in providing observability & Monitoring tooling, Change Automation tooling, security tool integration, automation support, and business continuity program for day-to-day operations.

Qualifications

  • BS or MS in Computer Science or closely related field.
  • 6+ years of work experience, with 4+ years in DevOps and Observability.
  • Certification in any cloud hosting platform such as Google/Azure/AWS/Oracle

Skills:

  • Hands-on Experience on Monitoring and Observability tooling development and Configuration.
  • Experienced in working with tools like Prometheus, Grafana, AppDynamics, Loki, Alertmanager etc.
  • Strong knowledge of Enterprise application architectures hosted on the cloud.
  • Knowledge of Infrastructure Automation Tools - Terraform, Chef, Ansible, Saltstack, Puppet, etc (Infrastructure As Code)
  • Knowledge of phases of the Software development life cycle
  • Knowledge on SW delivery pipeline, development environment, Build & Integration.
  • Comprehensive Knowledge of Continuous Integration Skills - Version control, builds, and remediation.
  • Strong automation skills (tool agnostic) and the ability to drive initiatives to automate processes
  • SW Testing - Unit, Integration, System, Load testing, performance, security, regression testing
  • Scripting & Language Skills - Python, Go, Groovy(/Java), Shell, Perl etc
  • Software Security Skills: Knowledge of PCI-DSS, HIPPA, SOX, GDPR, and CCPA Standards and Policies and the associated certification and audit processes
  • Knowledge of Integration of Software security tools to the CICD process.
  • Experience working with applications developed in Java, Node.js, Python, restful web services, and APIs.
  • Strong operational experience in Linux/Unix environment
  • Experience with REST-style web services / APIs.
  • Context-switch between multiple projects/codebases/concepts with ease
  • Understand software development at a fundamental level, use the best tools for the job, and always think about the future (at scale) when architecting solutions
  • Knowledge of fundamental aspects of release automation (packaging, dependencies, promotion, deployment, compliance)
  • Knowledge of the desired tools for release automation, Jenkins, repository management (SVN, GIT), and deploying software through scripts (ANT, Make, Shell script, Golang scripts)
  • Experience with technologies like Kafka, Docker, Elasticsearch, continuous integration (Drone, Jenkins, Travis, Bamboo) and understanding its benefits, workflows, etc.
  • Experience on project and ticket management tools such as JIRA and insight on quality analysis as well
  • Exposure in integrating testing tools such as Selenium, QTest Manager, etc
  • Proficient in some of these: Chef, GitHub, DevOps, Dockers, Jenkins, Black chair
  • Cloud experience (SaaS and PaaS) on Public Cloud e.g., AWS, Google, Azure, Oracle
  • Develop & implement DevOps Solutions in the areas of Observability, Change Automation, and Service Management.
  • Handles cross-functional collaboration to develop tools for secure, scalable, and reliable systems.
  • Identify, integrate, monitor, and improve infosec controls by understanding business processes.
  • Implements DevOps design principles and security best practices using relevant skills and experience.
  • Lead & influence the architectural design of features by determining quality & adhering to specifications.
  • Design & build solutions that move data from internal solutions to cloud-based solutions.
  • Responsible for Architecture/Design comprising, Security, Risk & Compliance
  • Monitor the delivery of solutions between architecture, time, cost, and quality and provide an Assessment of Costs and Benefits, and anchor continuous improvement.
  • Adopt the OCI standard tools and DevOps processes.
  • Tenets and best practices of Continuous Monitoring and Observability solutions.
  • Provides solutions for Continuous Delivery and Deployment (CD) and Change Management Automation
  • Provides solutions for Continuous Monitoring (CM) - monitoring and analysis of infrastructure, process, and applications
  • Provides solutions for Tooling for Checkout, Compile, Package, Verify, Quality Scan, Publish Artifact, Deploy to Development and Automation Testing
  • Provides solutions for Build & Release tools on Linux and Windows VMs/ Compute instances and containers.
  • Provides Code Quality Metrics, assists the team in improvement strategies
  • Performance Testing and tools Release processes
  • Ability to gel well with Agile teams in real spirit (culture & mindset)
  • Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation, and refinement.
  • Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning, and launch reviews.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Participate in incident handling and other related duties to support the information security function.
  • Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
  • Practice sustainable incident response and blameless postmortems.
  • Collaborate with Agile teams in defining technical requirements and best practices with containerized and cloud-native applications.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Developer in Santa Clara, CA vacancy
  •  ...with software, platform, and networking teams to improve service reliability and deployment workflowsDeploy and maintain network monitoring,...  ...in the on-call rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar roleHave... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  •  ...provisioning, upgrades, patching, and deletion.Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    17 hours ago
  • $168k - $270.25k

    NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization...  ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...activities such as system design consulting, developing software platforms and frameworks,...  ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $267k - $356k

     ...most demanding compute workloads in the industry, which means reliability and performance aren't just goals—they're the baseline. We're looking...  ...of software-defined storage across new and existing sites using tools such as Ansible, Jenkins etc.Work with hardware and... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $148k - $235.75k

     ...Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself (not the... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $101k - $161k

     ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  .../Pulumi, Bash. You will be expected to develop, operate, and work with many different...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-...  ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $98.9k - $228.7k

    What you can expectWe are hiring a Senior DevOps Engineer to ensure reliability, scalability, and operational excellence for our real-time communications platform. This platform supports audio/video conferencing, recording, and live-streaming functionalities. The position... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Zoom

    San Jose, CA
    2 days ago
  • $192.4k - $275.8k

     ...of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...— this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  •  ...Job Title: Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must...  ...meet FedRAMP High and IL5 requirements without sacrificing developer velocity Support ATO processes, including documentation... 
    Senior
    Full time

    SFE

    San Jose, CA
    2 days ago
  •  ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 8+ years of experience in SRE, DevOps, or platform engineering Hands-on experience... 
    Senior
    Full time

    SFE

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers...  ...Chips Simulation as a trusted and reliable virtual platform.What you will be doing:...  ...simulation environment interactionsImprove developer experience for teams using simulation infrastructure... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $307k - $427k

     ...to define the long-term roadmap for spatial intelligence.Mentor senior technical leads and staff engineers, fostering a culture of...  ...software, and hardware. Teams across this area research, design, and develop new technologies to make our user's interaction with computing... 
    Senior

    Google

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...our GPU Software team. In this role, you will help design and develop key components of our production GPU kernel drivers and embedded...  ...environmentsCollaborate with globally distributed teams to deliver scalable, reliable, and high-impact GPU software solutionsWhat we need to see: BS... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to the market! This role involves developing and productizing innovative solutions that will transform transportation and the field of self... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...motivated engineers who bring GeForce NOW to life! As a member of the GeForce Now Platform Engineering team, you will help design, develop, and optimize high performant GeForce Now gaming servers with state-of-the-art NVIDIA GPUs. Are you passionate about working in a cross... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...environment with fast pace and agility.What you'll be doing:Developing and triaging platform drivers which goes into SOCsBuilding sophisticated...  ...debugging is invaluableExperience working on system level reliability and resiliency features.Familiarity with system level... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $152k - $241.5k

    Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is an engineering...  ...facing GPU cloud services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    17 hours ago
  • $184k - $287.5k

     ...based infrastructure that allow agents to interact with internal developer systems (Gerrit, NVBugs, CI/CD, Perforce, Slack)Design and...  ...across GPU SW teams by shipping agentic services that are fast, reliable, and meaningfully better than the manual alternative!What we... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior System Software Engineer to work in our Tegra system software group. The best candidates will have excellent...  ...If you're a creative software engineer with a real passion for developing products with new technology, we want to hear from you. Join us... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...Proton titles.Identifying changes to API usage to improve performance and communicating via appropriate channels with third-party developers.Implementing driver performance improvements and resolving driver defects.Collaborating with engineers on the team and across... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...lower latency in batch ETL workloads. Apache Spark is the most popular data processing engine in data centers. At NVIDIA, we are developing an open source plugin to accelerate Spark applications on GPUs without any code changes.What you'll be doing:Enable C++ native execution... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...supercomputing, gaming, and visualization. As a Senior System Software Engineer on the NvSci...  ...and maintainability, and elevate developer experience.Evaluate trade-offs in resource...  ...generative AI technologies to improve software reliability, maintainability, and scalability.What... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $170k - $277k

     ...assisted coding tools (e.g., Cursor, Claude); you optimize your workflow and the team’s workflow around them.Strong desire to mentor and develop junior engineering talent, with a demonstrated ability to transition from a high-impact individual contributor to a technical... 
    Senior
    Full time
    Remote work
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...product lines and OEMs. We are looking for a highly motivated Senior Software Systems Engineer with a strong foundation in software...  ...hard-working attitude.Ways to Stand Out From The Crowd:Experience developing ADAS software.Deep understanding of real-time operating systems... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    17 hours ago
  • $184k - $287.5k

     ...and efficiency for our next generation of datacenter products, including CPUs and CPU+GPU Superchips. What you will be doing:Design, develop, test, and optimize software for our next-generation SoCs. In both pre-silicon and post-silicon phases of execution.Review... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA Solutions Engineering team is searching for engineers to help develop and bring NVIDIA’s autonomous vehicle platform to the world. You will work on state of the art technologies alongside experts in Deep Learning, Computer Vision, and vehicle control for NVIDIA’s... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...looking for energetic, enthusiastic and technologically savvy engineers to join the Tegra Core Firmware team. Here we architect and develop the boot stack firmware for the flagship Tegra chipset which are the core components for high-compute platforms for automotive,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Developer. Be the first to apply!