Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

$163.62k - $212.71k

iSpot

Job Description

Job Description

Immigration / Work Authorization Notice: Applicants must be currently authorized to work in the United States. iSpot is not able to sponsor or take over sponsorship of an employment visa for this position at this time.

iSpot competes for the best talent. Our compensation packages consist of salary and equity in one of Seattle's hottest start-ups, as well as other standard benefits. Most importantly, we provide a really interesting working experience, and the chance to contribute to the success of something great.

What You'll Be Part Of:

iSpot.tv is changing how brands, agencies, and networks measure and assess the impact of TV advertising. We deal with BIG data, operating mainly in AWS with multiple Kubernetes clusters and thousands of servers. We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for growth.

You will be a key member of our SRE leadership team, focused on empowering developers to build, test, and deploy applications faster and more efficiently. You will both lead the team and remain hands-on in designing, building, and maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software development lifecycle. Your work will directly impact developer happiness and the speed at which we can deliver innovative features to our customers.

Responsibilities:

We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and performance of our core production systems while significantly enhancing the internal developer experience. This role sits at the intersection of operations and development, requiring deep technical expertise, strong leadership, and a passion for optimizing the entire software development lifecycle (SDLC).

Our team consists of senior engineers who work together with minimal supervision to attain those goals. Candidates must possess deep operational experience with AWS and Kubernetes to support teams utilizing these systems. You will lead the technical direction of the team while remaining a key individual contributor. You will be responsible for creating a culture of engineering excellence, designing self-service platforms, and fostering alignment across all engineering teams to accelerate product delivery and maintain world-class service stability.The key responsibilities are:

  1. System Reliability and Operations (SRE Focus)
  • Platform Design and Management: Architect, build, and maintain scalable, highly available, and reliable cloud infrastructure in AWS leveraging modern container orchestration technologies.
  • Data Pipeline Reliability: Serve as the reliability and cost optimization expert for high-volume, data-intensive workloads. Focus on optimizing and ensuring the stability of distributed data processing engines, specifically Apache Spark and related ecosystems (e.g., EMR, Databricks, Glue).
  • Observability and Monitoring: Establish comprehensive observability practices by defining SLIs/SLOs, implementing advanced monitoring, alerting, and logging solutions to quickly identify and resolve system anomalies.
  • Automation: Drive automation across all operational aspects, including infrastructure provisioning (Terraform), scaling, deployment, and incident response, minimizing toil and manual effort.
  • Incident Management: Lead and participate in the incident response lifecycle, performing thorough post-mortems to derive actionable insights and implement preventative measures to improve system resilience.
  • AIOps: Define and champion the strategic roadmap for AI/ML integration within SRE, establishing organizational best practices for AIOps, automated incident remediation, Toil Reduction via LLMs, and Automated Root Cause Analysis (RCA) and the governance of LLM-driven tooling to enhance system observability and resilience.
  1. Developer Experience and Productivity (DevEx Focus)
  • Platform Strategy: Design, implement, and champion self-service tools, internal developer portals, and services that empower engineering teams to manage their infrastructure and deployments independently and efficiently.
  • AI Developer Tools: Lead the standardization of AI developer assistants by architecting and maintaining global 'steering files' and context-configuration standards, ensuring AI-generated code aligns with our specific patterns, security protocols, and architectural guardrails.
  • CI/CD Optimization: Own and continuously improve the CI/CD pipelines, reducing build times, streamlining deployment workflows, and integrating best practices for testing, security (Shift Left), and code quality. Maintain and improve our container orchestration and deployment tools, leveraging Kubernetes, Helm, and ArgoCD to create seamless developer workflows.
  • KPIs: Develop, implement, and maintain a set of key performance indicators (KPIs) to measure and improve the developer experience across all of Engineering.
  • Mentorship and Documentation: Guide and mentor senior engineers, promoting SRE/DevEx principles. Develop clear, comprehensive documentation and tutorials to ensure seamless adoption of new tools and platforms.
  • Cost and Efficiency: Strategically identify and implement opportunities for cloud cost optimization and resource efficiency without compromising reliability or performance.

III. Strategic Leadership and Cross-Team Alignment

  • Architecting the Roadmap: Define, champion, and communicate the long-term technical roadmap for the SRE and DevEx platforms, balancing immediate operational needs with strategic, future-state goals.
  • Driving Cross-Team Alignment: Act as a critical liaison between infrastructure, security, and product development teams. Proactively drive cross-team alignment on architectural standards, tooling choices, and development workflows to ensure consistency and shared accountability for system health.
  • Bottleneck Identification and Mitigation: Systematically identify engineering bottlenecks, friction points, and points of organizational toil within the SDLC. Implement targeted solutions—whether technical, process-based, or organizational—to mitigate these constraints and enhance overall engineering velocity.
  • Planning and Execution: Collaborate with engineering leadership to transform the strategic roadmap into actionable, prioritized plans, securing cross-functional buy-in and resources for successful execution.

Qualifications and Education Requirements:

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 10+ years of relevant experience in software engineering, cloud architecture, and/or Site Reliability Engineering, with at least 3 years in a leadership or lead contributor role.
  • Deep expertise of AWS, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA and S3.
  • Strong proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation).
  • Specialized experience in optimizing large-scale data platforms, specifically with Apache Spark. Proven ability to profile, troubleshoot, and tune Spark jobs for performance, cost, and reliability.
  • 5+ years of experience with Kubernetes and containerization in general, including associated tools (kubectl, Helm, ArgoCD).
  • Strong knowledge of AWS cost optimization.
  • TCP/IP networking, including routing and AWS security groups.
  • Excellent knowledge of CI/CD concepts and experience developing associated pipelines in CircleCI.
  • Proficient in high-level scripting languages, including shell scripting, Python, and/or JavaScript.
  • Experience with OTel and monitoring tools such as Splunk or DataDog. Experience with native AI observability tools is a plus.
  • Experience with evaluating and rolling out GenAI tools for improving developer efficiency.
  • Excellent communication, collaboration, and stakeholder management skills, with proven experience driving technical initiatives across multiple teams.
  • Experience with researching and selecting new/modern developer toolsets and assisting teams in adopting them including vendor assessments, security assessments and procurement process.
  • Experience in Ad-Tech or "BIG Data" processing organization is highly preferred

Target cash compensation range: $163,620 - $212,710 USD Annually

We are committed to providing competitive, market-informed compensation. The cash compensation above includes base salary, variable commission for employees in eligible roles, and annual bonus targets for eligible roles. In addition to cash compensation, all full time iSpotters are eligible to participate in iSpot's equity plan to receive stock options. Non-exempt roles will also be eligible for (pre-approved) overtime pay. Individual compensation packages are influenced by different factors unique to each candidate, including their skills, experience, qualifications and other job-related reasons.

For more information on total rewards package, go HERE

Hybrid & Flexible Workplace Policy

iSpot supports a hybrid and flexible workplace. Depending on location and work responsibilities, employees may be designated as full-time or part-time office-based or a fully remote employee. A hybrid work schedule indicates that you work in the office some days and work from home other days. The best hybrid workplaces allow for flexibility while also encouraging consistency.

Those local or living in surrounding areas to one of our offices (Bellevue, WA or New York, NY) will work a hybrid schedule, coming into their local office 1-3 days a week. While those in a role, not office-based and located further away from our offices, will work a fully remote schedule. If you have questions regarding exact details of our hybrid & flexible workplace policy, please let your recruiter know and they will discuss with you further.

#LI-Hybrid

If you don't feel you met every single requirement for the role, don't rule yourself out. Please apply anyway!

iSpot is an equal opportunity employer. All applicants will receive consideration for employment without regard to race, ethnicity, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. If you need assistance and/or a reasonable accommodation due to a disability during the application or the recruiting process, please contact our HR team.

California Residents applying for positions at iSpot can access our California Consumer Privacy Act here.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Bellevue, WA vacancy
  •  ...scale. This role is data-centric software engineering . We hold a high bar for quality: you’...  ...data lifecycle. What You’ll Do As a Principal Software Development Engineer, you bring...  ...paths. Institutionalize observability and reliability engineering for data: SLOs, monitoring,... 
    Principal
    Contract work

    The Consensus

    Bellevue, WA
    2 days ago
  •  ...care—about each other, about UiPath, and about our larger purpose. Could that be you? Your Mission We are looking for an engineering leader that operates as a Builder. Be immersive in shaping the 'what' and 'why' of products to build in addition to driving... 
    Principal
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    4 days ago
  • $274k - $376.2k

     ...This is a hybrid role. It requires going to the local office 3 times a week. The Auth0Lab Team We are a small team of engineers exploring new Auth0 products and features ideas. We take things from 0 to 1 and we look to shape the future of identity. Our team... 
    Principal
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Bellevue, WA
    1 day ago
  • $117.2k - $313.7k

     ...Job Category Software Engineering Job Details About Salesforce...  ...- Public Cloud (Senior/Lead/Principal) Note: By applying to the...  ...on our platform to be highly reliable, lightning fast, supremely secure...  ...experience balancing live-site management, feature delivery,... 
    Principal

    Salesforce.Com Inc

    Bellevue, WA
    19 hours ago
  • $164.59k - $246.79k

     ...Cities, Uncrewed Aircraft Systems (UAS), and Airspace Management including Urban Air Mobility (UTM). Echodyne is seeking a Principal Software Engineer, Cloud Platform, to join our fast‑growing team. Who You Are You are a high‑caliber technical leader who specializes in... 
    Principal
    Full time
    Temporary work
    Remote work
    Flexible hours

    Echodyne Corp

    Kirkland, WA
    2 days ago
  • $138.82k - $208.54k

     ...Description The Principal Software Developer uses a wide application of principles, theories, and technology concepts related to Human...  ..., or closely related. ~10 years experience as a Software Engineer or closely related occupation. ~10 years experience in application... 
    Principal
    Minimum wage
    Full time
    Work at office
    Local area
    Remote work
    Shift work

    Providence Service

    Bellevue, WA
    5 days ago
  • $227.8k - $338.8k

     ...language models, agentic frameworks, and production cloud engineering. We are hiring a Principal Engineer to be the technical anchor for this product....  ...tooling that customers build on. It has to be fast, reliable, and clear to work with. Lead technical execution across... 
    Principal
    Local area

    NetApp

    Bellevue, WA
    2 days ago
  • $197.3k - $313.7k

     ...you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI...  ...you are the future of Salesforce. Salesforce is seeking a Principal Software Engineer (PMTS) to join the Tableau Agent Analytics Agentforce... 
    Principal

    Salesforce.Com Inc

    Bellevue, WA
    1 day ago
  •  ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the Commercial and Investment Bank, you will solve complex and broad business... 

    J.P. Morgan

    Seattle, WA
    22 days ago
  • $280k - $330k

     ...scale. This role is data‑centric software engineering at a very high bar: you own substantial...  ...across the team. What You’ll Do As a Principal Software Development Engineer, you...  ...execute. Institutionalize observability and reliability engineering for data: SLIs/SLOs where... 
    Principal

    Auger Services Inc

    Bellevue, WA
    3 days ago
  • $304k

     ...Principal Engineer II At Snowflake, we are powering the era of the agentic enterprise. To usher...  ...: Design and implement highly reliable, multi-tenant system internals that handle...  ...the job posting on the Snowflake Careers Site for salary and benefits information: careers... 
    Principal
    Flexible hours

    Streamlit

    Bellevue, WA
    2 days ago
  •  ...of the world's most influential companies. As a Senior Principal Software Engineer at JPMorganChase within the CDAO AI/ML Data Platforms Team,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan,... 
    Principal

    J.P. Morgan

    Seattle, WA
    23 days ago
  • $117.2k - $313.7k

     ...Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM,...  ...all levels: Mid Level, Senior, Lead and Principal. *IN SCHOOL OR GRADUATED WITHIN THE LAST...  ...Benefits & Perks Check out our benefits site which explains our various benefits, including... 
    Principal
    Work at office
    Flexible hours
    3 days per week

    Salesforce.Com Inc

    Bellevue, WA
    2 days ago
  • $190k

     ...Principal DVM – Hillsboro, OR Ready to step into a role where you can shape the future of a practice, enjoy a loyal client base, and have a facility designed with veterinary workflows in mind? We are on the lookout for our next Principal DVM! Here’s the scoop:... 
    Principal
    Relocation package

    Peoplepack LLC

    Bellevue, WA
    1 day ago
  • $142.3k - $263.3k

     ...Senior Site Reliability Engineer Apple Services Engineering Cloud Service Infrastructure team is one of the most exciting examples of Apple...  ..., Operating Systems and Software ~ Understanding of SRE principals, including monitoring, alerting, error budgets, fault analysis... 
    Relocation

    Apple

    Seattle, WA
    4 days ago
  • $160k - $250k

     ...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able... 

    Hive

    Seattle, WA
    2 days ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    3 days ago
  •  ...people like you. As a Senior Principal Architect at JPMorganChase...  ...adoption of AI-assisted reliability workflows across SDLC/...  ...or certification on software engineering concepts and 10+ years applied...  ...comprehensive health care coverage, on-site health and wellness centers,... 
    Principal

    J.P. Morgan

    Seattle, WA
    12 days ago
  •  ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise... 
    Local area

    Comtech LLC

    Seattle, WA
    2 days ago
  • Job Description Job Description Part-Time Associate Veterinarian – (King County, WA) Join a newly remodeled practice that blends modern efficiency with top-tier care! We're looking for an experienced DVM 2–2.5 days per week to join our talented team. Why you'...
    Principal
    Part time
    2 days per week

    Peoplepack LLC

    Bellevue, WA
    19 days ago
  • $157.3k - $212.9k

     ...and free financial coaching. About the Role We are seeking a Principal Engineer to serve as a senior technical authority on enterprise...  ...increase delivery speed, maintainability, and operational reliability. Modernize existing automation capabilities by transitioning... 
    Principal
    Work experience placement
    Local area

    T-Mobile

    Bellevue, WA
    2 days ago
  • $276k - $414k

     ...addition to Bitmoji, Saturn, and other digital services. Snap Engineering teams build fun and technically sophisticated products that...  ...always execute with privacy at the forefront. We’re looking for a Principal Software Engineer to join Snap Inc! What you’ll do: Design,... 
    Principal
    Full time
    Temporary work
    Live in
    Work at office
    Local area

    Snap Inc.

    Bellevue, WA
    1 day ago
  •  ...to our people, and the incredible connections we get to make in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations team which is responsibility for observability and... 

    Kaav Inc.

    Seattle, WA
    1 day ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Temporary work

    Epic Games (Portuguese)

    Seattle, WA
    9 hours ago
  •  ...Client - Airbnb/Altimetrik ( Client remote / Hybrid from Altimetrik office ) Pay Rate -65 C2C Position: Site Reliability Engineer (SRE) - Backend/Cloud Must-Have Experience: Familiarity with cloud-hosted systems such as AWS or GCP . Hands... 
    Work at office
    Remote work

    Futran Tech Solutions Pvt. Ltd.

    Seattle, WA
    2 days ago
  • $100k - $170k

     ...who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll... 
    Flexible hours
    Shift work

    Nscale

    Seattle, WA
    3 days ago
  • Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...

    Syntricate Technologies

    Seattle, WA
    3 days ago
  •  ...that the residential, construction & building product industries operate across the globe. We are looking for a Manager, Site Reliability Engineering to be part of revolutionizing these industries. What You Will Do Lead and grow a team of site reliability engineers.... 

    Paradigm

    Seattle, WA
    2 days ago
  • $276k - $414k

     ...We’re looking for a Principal Software Engineer to join the Ads Platform team at Snap. What you’ll do: Build the next generation ads formats and backend infrastructure and solutions to deliver more clicks and conversion and value for advertisers. Work end-to-end across... 
    Principal

    Snap

    Seattle, WA
    2 days ago
  •  ...operations, you’ll do your best work here. The Role As a Principal Platform Engineer at Gradial, you will shape the foundation our platform runs...  ...leverage, and the opportunity to define how platform reliability looks at an AI‑native company. What You’ll Own Own the reliability... 
    Principal

    Outsiders Fund

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!