Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

OneStream Software LLC

Role Description

As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If you enjoy staying at the forefront of technology and automating infrastructure deployments, then this is the job for you. This vital role within Cloud Services requires knowledge and experience designing, implementing, and monitoring scalable and secure cloud services. The employee is expected to work well in a small team and willing to share responsibilities with other team members as needed. You will interact with internal staff, managers, and customers to implement and maintain operations. A passion for technology and learning, and the ability to grow others are vital for success in this role.

Primary Duties and Responsibilities

  • Implement application/infrastructure observability solutions to ensure desired application availability, reliability, and performance.
  • Participate in regular On-Call rotations and share details related to incidents and their resolution through post-mortem reports and regular review meetings.
  • Proactively partner with Product and Engineering teams to identify, develop, deploy, and maintain reliable systems and services.
  • Influence and create new designs, architectures, standards, and methods for large-scale systems.
  • Sustain a high level of reliability for key services and automated systems.
  • Automate processes to improve reliability, performance, and availability.
  • Update technical documentation, workflows, and knowledge base articles.
  • Provide feedback in pull requests and peer coding reviews.
  • Implement codified automated solutions that build integrations between Dynatrace, Azure DevOps and Jira.
  • Solid knowledge in focused areas of OneStream Software.
  • Ability to mentor others in several technical areas.
  • Understanding practical use of SOC/FedRAMP controls to assist Compliance and Security teams.

Qualifications

  • BS/BA in computer science, engineering, or technology-related field (or equivalent work experience).
  • Proven work experience as a Site Reliability Engineer or in a similar role.
  • 6+ years of cloud infrastructure and software development experience.
  • 2+ years hands-on experience of Azure Kubernetes Services (AKS) with container-based deployment skills or other platforms such as OpenShift, GKS, EKS.
  • Advanced understanding of APM and observability tools such as Dynatrace, AppInsights, DataDog, Log Analytics, New Relic, Prometheus and Grafana.
  • Advanced understanding of Infrastructure-as-Code (IaC) concepts and tooling (Terraform, CloudFormation templates, Bicep or ARM templates) on Microsoft Azure, Amazon Web Services (AWS), or Google Cloud Platform (GCP).
  • Deep knowledge of Configuration Management/Orchestration utilities such as Ansible, PowerShell DSC, Chef, and Puppet.
  • Advanced understanding of cloud concepts including elasticity, security, and identity management.
  • Well versed familiarity with Agile Development methodologies utilizing Jira or Azure DevOps Boards.
  • 6+ years of hands-on experience with the following technologies, tools, and concepts:
    • Automating processes using PowerShell, Bash, CLI, REST APIs, python, ARM Templates or other scripting languages.
    • Comfortable leveraging source control tools such as Git, Azure DevOps, or GitHub.
    • Knowledge of container orchestration platforms such as Kubernetes, OpenShift, AKS, GKS or helm.
    • Microsoft Azure, Amazon Web Services (AWS) or Google Cloud (GCP).

Preferred Education and Experience

  • Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider.
  • 6+ years of relevant Azure experience deploying and managing leveraging Infrastructure-as-Code (IAC) concepts.
  • Experience with Microsoft and .NET (.NET, C#, SQL).
  • Experience writing efficient and reliable code in a development environment.
  • Debian, Ubuntu, Alpine or other distributions of the Linux operating systems.
  • Deep knowledge and understanding of containerized applications, with special attention to reliability and monitoring of those containerized applications.

Knowledge, Skills, and Abilities

  • Deal well with ambiguous/undefined problems.
  • Ability to self-motivate and work independently.
  • Strong organizational and prioritization skills.
  • Ability to find and apply effective solutions to emerging problems and challenges.
  • Strong attention to detail.
  • Comfortable communicating with all levels of management and engineering.
  • Ability to get up to speed quickly with modern technologies and services.
  • Ability to multitask on a variety of projects.

Travel

  • Travel Requirement: Travel is not expected to exceed 5%.

Benefits

  • Excellent Medical Plan.
  • Dental & Vision Insurance.
  • Life Insurance.
  • Short & Long Term Disability.
  • Vacation Time.
  • Paid Holidays.
  • Professional Development.
  • Retirement Plan.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Remote vacancy
  • $140k - $150k

    WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Weekend work

    National Basketball Association

    Secaucus, NJ
    3 days ago
  • $87.1k - $157.45k

     ...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across...  ...to come in and help us build systems that stay reliable when things get complicated.We need a Site Reliability Engineer who has experience building, deploying... 
    Suggested
    Full time
    Work from home
    Flexible hours

    Leidos

    Chantilly, Loudoun County, VA
    2 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    3 days ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    2 days ago
  • Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance, reliability, and cost-effectiveness of the cloud-based infrastructure and related systems supporting... 
    Suggested
    Full time
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Internship
    Local area
    Worldwide

    Life.Church

    Edmond, OK
    2 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    2 days ago
  •  ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology...  ...AI-driven solutions that accelerate development and improve reliability. Your work will directly influence how GM Financial leverages... 
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    2 days per week

    GM Financial

    Arlington, TX
    2 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Remote work

    Patterson-UTI

    Houston, TX
    5 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    2 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $86.9k - $198k

    Site Reliability Engineer, SeniorThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    4 days ago
  •  ...some of the most important challenges in global education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work towards becoming a Subject Matter Expert in the cloud space... 
    Remote work

    Intelliswift

    Durham, NC
    2 days ago
  •  ...: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Remote work

    SRI Tech

    Plano, TX
    2 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    McLean, VA
    4 days ago
  • $165k - $190k

     ...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-...  ...security platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  • $35 - $44 per hour

    DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and... 
    Remote work

    KForce

    Atlanta, GA
    5 days ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    3 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    Govcio

    Arlington, VA
    2 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    2 days ago
  •  ...and responsible for ensuring the availability, scalability, and reliability of systems and applications.What will be your responsibilities...  ...using tools like Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-to-date with industry trends... 
    Work at office
    Remote work

    Interactive Brokers

    Greenwich, CT
    6 days ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are... 
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...alert behavior preferred Exposure to reliability engineering concepts such as SLOs/SLIs and...  ...office#LI-KC1#GMFjobsAbout The Role:The Site Reliability Engineer under the general... 
    Work experience placement
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    Shift work
    2 days per week

    GM Financial

    Arlington, TX
    5 days ago
  • $15k

     ...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills... 
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    5 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $118.6k - $195.68k

    Job SummaryThe Red Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our Red Hat Hybrid OpenShift Platforms (on-prem & cloud). As a Senior Engineer, you will contribute to running Red Hat OpenShift at scale... 
    Permanent employment
    Full time
    Contract work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Red Hat

    Raleigh, NC
    5 days ago
  • $130k - $180k

     ...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails... 
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours

    Imanage

    Chicago, IL
    4 days ago
  • $102.1k - $202.2k

     ...per yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole type...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...workloads. As a Site Reliability Engineer II, you will take ownership of reliability... 
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    6 days ago
  • $105.6k - $145.2k

    Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience supporting multiple connected Cloud-based products? Trimble is a global technology... 
    Ongoing contract
    Full time
    Work at office
    Local area
    Worldwide

    Trimble Navigation

    Westminster, CO
    3 days ago
  • $108.08k - $172.5k

    Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts and applications to maintain services. Monitor metrics, system health and analyze reports. Provide on-call rotation support for production... 
    Full time
    Remote work
    Worldwide

    CME- Group

    Chicago, IL
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!