Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

$163.62k - $212.71k

iSpot.tv

Immigration / Work Authorization Notice:Applicants must be currently authorized to work in the United States. iSpot is not able to sponsor or take over sponsorship of an employment visa for this position at this time.iSpot competes for the best talent. Our compensation packages consist of salary and equity in one of Seattle’s hottest start-ups, as well as other standard benefits. Most importantly, we provide a really interesting working experience, and the chance to contribute to the success of something great.What You’ll Be Part Of:iSpot.tv is changing how brands, agencies, and networks measure and assess the impact of TV advertising. We deal with BIG data, operating mainly in AWS with multiple Kubernetes clusters and thousands of servers. We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for growth.You will be a key member of our SRE leadership team, focused on empowering developers to build, test, and deploy applications faster and more efficiently. You will both lead the team and remain hands-on in designing, building, and maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software development lifecycle. Your work will directly impact developer happiness and the speed at which we can deliver innovative features to our customers.Responsibilities:We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and performance of our core production systems while significantly enhancing the internal developer experience. This role sits at the intersection of operations and development, requiring deep technical expertise, strong leadership, and a passion for optimizing the entire software development lifecycle (SDLC).Our team consists of senior engineers who work together with minimal supervision to attain those goals. Candidates must possess deep operational experience with AWS and Kubernetes to support teams utilizing these systems. You will lead the technical direction of the team while remaining a key individual contributor. You will be responsible for creating a culture of engineering excellence, designing self-service platforms, and fostering alignment across all engineering teams to accelerate product delivery and maintain world-class service stability.The key responsibilities are:System Reliability and Operations (SRE Focus)Platform Design and Management: Architect, build, and maintain scalable, highly available, and reliable cloud infrastructure in AWS leveraging modern container orchestration technologies.Data Pipeline Reliability: Serve as the reliability and cost optimization expert for high-volume, data-intensive workloads. Focus on optimizing and ensuring the stability of distributed data processing engines, specifically Apache Spark and related ecosystems (e.g., EMR, Databricks, Glue).Observability and Monitoring: Establish comprehensive observability practices by defining SLIs/SLOs, implementing advanced monitoring, alerting, and logging solutions to quickly identify and resolve system anomalies.Automation: Drive automation across all operational aspects, including infrastructure provisioning (Terraform), scaling, deployment, and incident response, minimizing toil and manual effort.Incident Management: Lead and participate in the incident response lifecycle, performing thorough post-mortems to derive actionable insights and implement preventative measures to improve system resilience.AIOps: Define and champion the strategic roadmap for AI/ML integration within SRE, establishing organizational best practices for AIOps, automated incident remediation, Toil Reduction via LLMs, and Automated Root Cause Analysis (RCA)and the governance of LLM-driven tooling to enhance system observability and resilience.Developer Experience and Productivity (DevEx Focus)Platform Strategy: Design, implement, and champion self-service tools, internal developer portals, and services that empower engineering teams to manage their infrastructure and deployments independently and efficiently.AI Developer Tools: Lead the standardization of AI developer assistants by architecting and maintaining global 'steering files' and context-configuration standards, ensuring AI-generated code aligns with our specific patterns, security protocols, and architectural guardrails.CI/CD Optimization: Own and continuously improve the CI/CD pipelines, reducing build times, streamlining deployment workflows, and integrating best practices for testing, security (Shift Left), and code quality. Maintain and improve our container orchestration and deployment tools, leveraging Kubernetes, Helm, and ArgoCD to create seamless developer workflows.KPIs: Develop, implement, and maintain a set of key performance indicators (KPIs) to measure and improve the developer experience across all of Engineering.Mentorship and Documentation: Guide and mentor senior engineers, promoting SRE/DevEx principles. Develop clear, comprehensive documentation and tutorials to ensure seamless adoption of new tools and platforms.Cost and Efficiency: Strategically identify and implement opportunities for cloud cost optimization and resource efficiency without compromising reliability or performance.III. Strategic Leadership and Cross-Team AlignmentArchitecting the Roadmap: Define, champion, and communicate the long-term technical roadmap for the SRE and DevEx platforms, balancing immediate operational needs with strategic, future-state goals.Driving Cross-Team Alignment: Act as a critical liaison between infrastructure, security, and product development teams. Proactively drive cross-team alignment on architectural standards, tooling choices, and development workflows to ensure consistency and shared accountability for system health.Bottleneck Identification and Mitigation: Systematically identify engineering bottlenecks, friction points, and points of organizational toil within the SDLC. Implement targeted solutions—whether technical, process-based, or organizational—to mitigate these constraints and enhance overall engineering velocity.Planning and Execution: Collaborate with engineering leadership to transform the strategic roadmap into actionable, prioritized plans, securing cross-functional buy-in and resources for successful execution. Qualifications and Education Requirements:Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.10+ years of relevant experience in software engineering, cloud architecture, and/or Site Reliability Engineering, with at least 3 years in a leadership or lead contributor role.Deep expertise of AWS, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA and S3.Strong proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation).Specialized experience in optimizing large-scale data platforms, specifically with Apache Spark. Proven ability to profile, troubleshoot, and tune Spark jobs for performance, cost, and reliability.5+ years of experience with Kubernetes and containerization in general, including associated tools (kubectl, Helm, ArgoCD).Strong knowledge of AWS cost optimization.TCP/IP networking, including routing and AWS security groups.Excellent knowledge of CI/CD concepts and experience developing associated pipelines in CircleCI.Proficient in high-level scripting languages, including shell scripting, Python, and/or JavaScript.Experience with OTel and monitoring tools such as Splunk or DataDog. Experience with native AI observability tools is a plus.Experience with evaluating and rolling out GenAI tools for improving developer efficiency.Excellent communication, collaboration, and stakeholder management skills, with proven experience driving technical initiatives across multiple teams.Experience with researching and selecting new/modern developer toolsets and assisting teams in adopting them including vendor assessments, security assessments and procurement process.Experience in Ad-Tech or “BIG Data” processing organization is highly preferredTarget cash compensation range: $163,620 - $212,710 USD AnnuallyWe are committed to providing competitive, market-informed compensation. The cash compensation above includes base salary, variable commission for employees in eligible roles, and annual bonus targets for eligible roles. In addition to cash compensation, all full time iSpotters are eligible to participate in iSpot’s equity plan to receive stock options. Non-exempt roles will also be eligible for (pre-approved) overtime pay. Individual compensation packages are influenced by different factors unique to each candidate, including their skills, experience, qualifications and other job-related reasons.For more information on total rewards package, goHEREHybrid & Flexible Workplace PolicyiSpot supports a hybrid and flexible workplace. Depending on location and work responsibilities, employees may be designated as full-time or part-time office-based or a fully remote employee. A hybrid work schedule indicates that you work in the office some days and work from home other days. The best hybrid workplaces allow for flexibility while also encouraging consistency. Those local or living in surrounding areas to one of our offices (Bellevue, WA or New York, NY) will work a hybrid schedule, coming into their local office 1-3 days a week. While those in a role, not office-based and located further away from our offices, will work a fully remote schedule. If you have questions regarding exact details of our hybrid & flexible workplace policy, please let your recruiter know and they will discuss with you further.#LI-HybridIf you don't feel you met every single requirement for the role, don't rule yourself out. Please apply anyway!iSpot is an equal opportunity employer. All applicants will receive consideration for employment without regard to race, ethnicity, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. If you need assistance and/or a reasonable accommodation due to a disability during the application or the recruiting process, please contact our HR team.California Residents applying for positions at iSpot can access our California Consumer Privacy Act here.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Bellevue, WA vacancy
  • $165k - $225.6k

     ...core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. THE SENIOR SITE RELIABILITY ENGINEER OPPORTUNITY Reporting to the Manager, Site Reliability Engineering, this role will help... 
    Suggested
    Permanent employment
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  •  ...Principal Site Reliability Engineer (IC4) As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure... 
    Principal

    Oracle

    Seattle, WA
    4 days ago
  • $200k - $350k

     ...Hybrid role: 3 days/week in office Role Overview As a Principal Software Engineer at MediaAlpha, you will be a technical leader accountable...  ...maintains standards for maintainability, scalability, security, reliability, privacy, and compliance at ecosystem scale Partner with... 
    Principal
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours
    3 days per week

    MediaAlpha

    Bellevue, WA
    1 day ago
  • Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
    Suggested

    The Dignify Solutions, LLC

    Bellevue, WA
    5 days ago
  • $174k - $253k

     ...Minimum qualifications Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5...  ...s degree in Computer Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when you treat operations... 
    Suggested
    Temporary work

    Google

    Kirkland, WA
    2 days ago
  • $194k - $267k

     ...something more than once, automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Bellevue, WA
    3 days ago
  •  ...scale. This role is data-centric software engineering . We hold a high bar for quality: you’...  ...data lifecycle. What You’ll Do As a Principal Software Development Engineer, you bring...  ...paths. Institutionalize observability and reliability engineering for data: SLOs, monitoring,... 
    Principal
    Contract work

    The Consensus

    Bellevue, WA
    1 day ago
  • $163.62k - $212.71k

     ...to contribute to the success of something great. As a Principal Software Development Engineer, you will be a hands-on leader and strategic thinker,...  ...measurement and data processing platform, ensuring scalability, reliability, and high performance. Lead the design of core... 
    Principal
    Full time
    Work experience placement

    PVH (Tommy Hilfiger/Calvin Klein)

    Bellevue, WA
    1 day ago
  • At Oracle Cloud Infrastructure (OCI), we are building the next generation of cloud services for autonomous and intelligent robotics. The OCI Robotics Cloud team is developing a secure, scalable, and cloud-native robotics platform that enables customers to seamlessly ...
    Principal
    Full time
    Flexible hours

    Oracle

    Seattle, WA
    1 day ago
  • $222.9k - $295.2k

     ...Your Impact We're looking for a Principal Software Engineer who enjoys solving difficult technical...  ...search and analytics systems that are reliable, scalable, and easy to evolve. Partner...  .... Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Principal
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    Kirkland, WA
    5 days ago
  • $102.1k - $202.2k

     ...per yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole type...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...workloads. As a Site Reliability Engineer II, you will take ownership of reliability... 
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    1 day ago
  •  ...platform, the backbone of our real-time insights, intelligent applications, and advanced decision-making capabilities. As a Principal Engineer for the Auger Platform, you will bring deep expertise in the following: What You'll Do Own end-to-end systems architecture across... 
    Principal

    Auger Inc.

    Bellevue, WA
    3 days ago
  • $280k - $330k

     ...scale. This role is data‑centric software engineering at a very high bar: you own substantial...  ...across the team. What You’ll Do As a Principal Software Development Engineer, you...  ...execute. Institutionalize observability and reliability engineering for data: SLIs/SLOs where... 
    Principal

    Auger Services Inc

    Bellevue, WA
    2 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter. Your Impact As a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    2 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type:...  ...Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...demanding workloads.We are seeking a Principal Site Reliability Engineering Manager to lead a team responsible... 
    Principal
    Ongoing contract
    Temporary work
    Fixed term contract
    Local area
    Immediate start
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago
  •  ...looking for people like you. As a Senior Principal Architect at JPMorganChase within...  ...reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain...  ...training or certification on software engineering concepts and 10+ years applied experience... 
    Principal

    JPMorgan Chase & Co.

    Seattle, WA
    1 day ago
  •  ...Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise content... 

    Comtech LLC

    Seattle, WA
    1 day ago
  •  ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER — HPC & AUTOMATION (SILICON ENGINEERING)At SpaceX we’re leveraging our experience in building rockets and spacecraft to... 
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    5 hours ago
  •  ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up to big things. The practice and philosophy of yoga informs our overall purpose to elevate the world through the power of practice. We are proud to be a growing global... 

    Samprasoft

    Seattle, WA
    2 days ago
  •  ...manage AI resources on Microsoft Azure, including AI Foundry and RAG solutions Monitor and ensure service uptime, availability, reliability, and latency Track and integrate SRE metrics with enterprise monitoring systems Support CI/CD and DevOps workflows using... 

    Tech M USAAvance Consulting

    Redmond, WA
    4 days ago
  •  ...Oracle in Seattle is seeking engineers to build large-scale distributed infrastructure for the cloud. The Object Storage Service team...  ...scalable systems, implement automation with IaC, and help ensure reliability and security in a multi-tenant environment. This role offers... 
    Principal

    Oracle

    Seattle, WA
    1 day ago
  • $264.1k - $369.74k

    Sr Principal Software Engineering - Enterprise Technology page is loaded## Sr Principal Software Engineering - Enterprise Technologylocations: Greater...  ...for:WA applicants is $264,103.00 - $369,743.85**Other site ranges may differ****Culture Statement****Export Control Regulations... 
    Principal
    Permanent employment
    Temporary work
    Local area
    Relocation

    Blue Origin LLC

    Seattle, WA
    1 day ago
  •  ...The Role Join us in revolutionizing the lending landscape. SoFi is seeking enthusiastic Principal Software Engineers who are ready to lead the technical and strategic evolution of our financial services platform in support of our goals that put our members in control... 
    Principal
    Full time
    Temporary work
    Work experience placement

    Sofi

    Seattle, WA
    9 hours ago
  •  ...Overview Tech Talent Specialist | USA | Around the whole SDLC Senior / Principal Software Engineer Some roles keep you fixing features no one notices. Others let you build tech that actually stops real cyber threats before they cause chaos. (This role is the second one... 
    Principal
    Home office

    Tact.ai

    Seattle, WA
    1 day ago
  • $94.3k - $156.9k

     ...and innovative future is now.   PSE's Financial Planning & Analysis team is looking for qualified candidates to fill an open  Principal Financial Analyst (Corporate FP&A) position! Specific details regarding the work arrangements for this position will be discussed... 
    Principal
    Temporary work
    Local area
    Flexible hours

    Puget Sound Energy

    Bellevue, WA
    2 hours ago
  • $196k - $294k

     ...ABOUT THE TEAM ~ We are seeking a Principal Software Engineer to lead architectural and strategic initiatives for our ArsenalOS Forge platform. You'll be pivotal in evolving Forge's capabilities, driving integration and scalability, and maturing its architecture for... 
    Principal
    Full time
    Work experience placement
    Local area
    Relocation package

    Anduril Industries

    Seattle, WA
    9 hours ago
  • $200k

     ...systems, focusing on services that ensure Synthesia is secure, reliable, and scalable for our largest customers. You will...  ...clear steps that can be delivered and validated iteratively. Engineers within Synthesia are empowered to contribute heavily to product... 
    Principal
    Full time
    For contractors
    Local area
    Visa sponsorship

    Synthesia Limited

    Seattle, WA
    1 day ago
  • $269.18k - $376.84k

     ...worldwide. About the Role As a Principal BSP Embedded Software Engineer, you will be a technical leader...  ...silicon to boot, initialize, and operate reliably in the demanding environment of...  ...temporary remote work exception while our sites are developed. We are looking... 
    Principal
    Permanent employment
    Temporary work
    Local area
    Remote work
    Worldwide

    Socket.dev

    Seattle, WA
    2 days ago
  • $194k - $267k

     ...this mission. If you are too, let's talk. Position Overview: We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    28 days ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!