Principal Site Reliability Engineer
Fidelity Investments
Job Description:Note: Fidelity will not provide immigration sponsorship for this position.Position Description:Delivers services at high scale and high availability with resilience by using automation and Infrastructure Code. Builds reliability into the ecosystem by applying best practices in resiliency engineering and observability by developing resiliency tools and capabilities for observability and chaos testing of pipelines. Combines systems and software engineering techniques with site reliability engineering practices to create reliable user experiences to support workplace investing, healthcare and defined benefits organizations. Assists teams scale through production insights, operational automation, developer guidance, real-time metrics, and automation. Provides training, support, and alignment to ensures Site Reliability Engineers have the skills, tools, and opportunities to accomplish engineering reliable systems. Provides Product and Platform teams with engineering expertise that enables them to clarify and achieve their system reliability goals. Partners with our key stakeholders in defining and adopting policies, processes and practices that lead to reliable Information Technology systems and measures compliance with those policies.Primary Responsibilities:Provides cloud support and enhances cloud capabilities according to Site Reliability Engineering principles -- observability, automation, and resiliency.Develops and enhances internal chaos framework to streamline chaos executions and reporting.Facilitates the adoption of chaos engineering by application teams, helping to perform chaos testing, and to analyze business-critical applications to understand the weaknesses and increase application resiliency.Develops and designs products created around the Site Reliability Engineering domain to improve stability and improve platform availability.Collaborates with business and technology teams to scale the products and automation across business units.Minimizes the impact of operational problems by developing strategies and tools to remediate the issues.Provides technical leadership on chaos testing for cloud and on-premises based applications.Develops scripts and applications to automate repeatable business processes.Advises senior management on technical strategy and tools.Mentors team members to build core competencies required in Site Reliability Engineering space.Education and Experience:Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) designing and automating container and Cloud-based platform products and infrastructure solutions within a production environment.Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology, Information Systems or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) designing and automating container and Cloud-based platform products and infrastructure solutions within a production environment.Skills and Knowledge:Candidate must also possess:Demonstrated Expertise (“DE”) developing and designing products created around Site Reliability Engineering pillars to improve stability and improve platform availability for containerized workloads and on-premises services using Kubernetes; managing and interpreting datasets using query languages; and developing reports using PowerBi and Grafana.DE managing Cloud and on-premises systems using infrastructure-as-code tools -- Azure ARM and Terraform; and building, operating, monitoring, logging, and alerting services of distributed systems at scale and utilizing modern monitoring tools – Datadog and Splunk.DE creating solutions that support DevOps practice for delivery and operations of services using Jenkins and Azure DevOps, Team Foundation Version Control, and Cloud Formation Template.DE maintaining scalability and resiliency of applications deployed on Amazon Web Services (AWS) and Azure using LAMBDA, API Gateway, Fault Injection Service (FIS), and Azure Chaos Studio; and developing scripts and applications to automate repeatable business processes and providing solutions to program needs for applications hosted on Windows and Linux using Python.#PE1M2#LI-DNIFidelity’s Onsite Working ModelFidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.Certifications:Category:Information TechnologyPlease be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.SummaryLocation: Westlake, TXType: Full time
$198.24k - $272.58k
We’re looking for a Principal Site Reliability Engineer to join Procore’s Compute Division to work on our FedRAMP initiative. In this role, you’ll help build Procore’s next-generation construction compute platform for others to build upon, including Procore developers,...PrincipalFull timeContract workWork at officeLocal areaImmediate start- ...You will provide cloud operations for Oracle National Security Realms. Responsibilities Escalation points for junior site reliability engineers during complex or high-impact incidents. Manage and execute complex manual Change Management tickets, by working...PrincipalTemporary workWork experience placementFlexible hoursNight shift
- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...PrincipalFull timeWork at office
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the...SuggestedFull timeWork at office
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SuggestedTemporary workCasual workWorldwide
- ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve... ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (...Full timeLocal area
$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...plays a pivotal role in engineering the reliable, globally connected, multi-cloud network... ...are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...Local areaRemote workWorldwideFlexible hours- ...in Ausin, TX** Our Opportunity: We are looking for a skilled engineer with disciplines that incorporate aspects of software systems... ...applications — including AI/ML-driven approaches to observability and reliability. What you’ll do: • Evangelize SRE mindset and solve problems...
$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to deliver...Full timeFlexible hours
- ...Job Description:Aboutthe Role:We arelooking for a Senior SRE to join our Platform Engineering team whereyou’llown the reliability, scalability, and operational excellenceof our workflow orchestration platforms – primarily Apache Airflow and BroadcomAutomic/UC4.This is...Full time
$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...- ...selected candidate for this role to work on site in the specified location(s).The Client... ...team is responsible for ensuring the reliability, scalability, and operational excellence... ...around the clock. As a Site Reliability Engineer, you will partner across application engineering...Full timeWork at office
- ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help... ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-...Full time
$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...Work experience placementWork at office- ...United States of America / Alberta / British ColumbiaTechnology – Engineering /Full-time - Permanent /RemoteAbout MegaportWe’re not your... ...goals are met.What You Will Be DoingImproving production reliability and system resilience within an SRE scoped teamChampioning high...Permanent employmentFull timeRemote workFlexible hours
$172k - $300k
Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software.As one of our founding SREs, you...Full timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...Work at officeLocal area
- Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This...Full time
- ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States (US).Position OverviewWe are seeking a highly skilled Site...Temporary workWork at officeRemote workFlexible hours
$127k - $249k
...MongoDB, Inc. is seeking an experienced Senior or Staff Engineer for their SRE, InfraSec team, responsible for guiding the security of cloud-based infrastructure. The role involves hands-on technical work and mentorship of a small team while collaborating with engineering...Remote workFlexible hours- ...Role: Site Reliability Engineer Location: Southlake / Austin, TX - Onsite 4 days weekly Duration: 12 Months Job Summary We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure,...For contractors
- ...Site Reliability Engineer Department: Infrastructure Employment Type: Full Time Location: Austin Reporting To: SRE Manager Description Are you an SRE ready to grow your impact at the world's largest AI cloud-native physical security company? Following...Full timeWorldwide
$140k - $215k
...processing trillions of events per day. As a Principal SRE, you will operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational... ..., while embedding directly with product engineering teams and their leadership to drive reliability...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- senior chief engineer Austin, TX
- general engineer Austin, TX
- project engineer assistant project manager Austin, TX
- chief design engineer Austin, TX
- principal infrastructure engineer Austin, TX
- principal cloud engineer Austin, TX
- chief engineer Austin, TX
- principal developer Austin, TX
- senior principal engineer Austin, TX
- engineering director Austin, TX


