Principal Site Reliability Engineer
$163.62k - $212.71kiSpot
Job Description
Job Description
Immigration / Work Authorization Notice: Applicants must be currently authorized to work in the United States. iSpot is not able to sponsor or take over sponsorship of an employment visa for this position at this time.
iSpot competes for the best talent. Our compensation packages consist of salary and equity in one of Seattle's hottest start-ups, as well as other standard benefits. Most importantly, we provide a really interesting working experience, and the chance to contribute to the success of something great.
What You'll Be Part Of:iSpot.tv is changing how brands, agencies, and networks measure and assess the impact of TV advertising. We deal with BIG data, operating mainly in AWS with multiple Kubernetes clusters and thousands of servers. We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for growth.
You will be a key member of our SRE leadership team, focused on empowering developers to build, test, and deploy applications faster and more efficiently. You will both lead the team and remain hands-on in designing, building, and maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software development lifecycle. Your work will directly impact developer happiness and the speed at which we can deliver innovative features to our customers.
Responsibilities:We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and performance of our core production systems while significantly enhancing the internal developer experience. This role sits at the intersection of operations and development, requiring deep technical expertise, strong leadership, and a passion for optimizing the entire software development lifecycle (SDLC).
Our team consists of senior engineers who work together with minimal supervision to attain those goals. Candidates must possess deep operational experience with AWS and Kubernetes to support teams utilizing these systems. You will lead the technical direction of the team while remaining a key individual contributor. You will be responsible for creating a culture of engineering excellence, designing self-service platforms, and fostering alignment across all engineering teams to accelerate product delivery and maintain world-class service stability.The key responsibilities are:
- System Reliability and Operations (SRE Focus)
- Platform Design and Management: Architect, build, and maintain scalable, highly available, and reliable cloud infrastructure in AWS leveraging modern container orchestration technologies.
- Data Pipeline Reliability: Serve as the reliability and cost optimization expert for high-volume, data-intensive workloads. Focus on optimizing and ensuring the stability of distributed data processing engines, specifically Apache Spark and related ecosystems (e.g., EMR, Databricks, Glue).
- Observability and Monitoring: Establish comprehensive observability practices by defining SLIs/SLOs, implementing advanced monitoring, alerting, and logging solutions to quickly identify and resolve system anomalies.
- Automation: Drive automation across all operational aspects, including infrastructure provisioning (Terraform), scaling, deployment, and incident response, minimizing toil and manual effort.
- Incident Management: Lead and participate in the incident response lifecycle, performing thorough post-mortems to derive actionable insights and implement preventative measures to improve system resilience.
- AIOps: Define and champion the strategic roadmap for AI/ML integration within SRE, establishing organizational best practices for AIOps, automated incident remediation, Toil Reduction via LLMs, and Automated Root Cause Analysis (RCA) and the governance of LLM-driven tooling to enhance system observability and resilience.
- Developer Experience and Productivity (DevEx Focus)
- Platform Strategy: Design, implement, and champion self-service tools, internal developer portals, and services that empower engineering teams to manage their infrastructure and deployments independently and efficiently.
- AI Developer Tools: Lead the standardization of AI developer assistants by architecting and maintaining global 'steering files' and context-configuration standards, ensuring AI-generated code aligns with our specific patterns, security protocols, and architectural guardrails.
- CI/CD Optimization: Own and continuously improve the CI/CD pipelines, reducing build times, streamlining deployment workflows, and integrating best practices for testing, security (Shift Left), and code quality. Maintain and improve our container orchestration and deployment tools, leveraging Kubernetes, Helm, and ArgoCD to create seamless developer workflows.
- KPIs: Develop, implement, and maintain a set of key performance indicators (KPIs) to measure and improve the developer experience across all of Engineering.
- Mentorship and Documentation: Guide and mentor senior engineers, promoting SRE/DevEx principles. Develop clear, comprehensive documentation and tutorials to ensure seamless adoption of new tools and platforms.
- Cost and Efficiency: Strategically identify and implement opportunities for cloud cost optimization and resource efficiency without compromising reliability or performance.
III. Strategic Leadership and Cross-Team Alignment
- Architecting the Roadmap: Define, champion, and communicate the long-term technical roadmap for the SRE and DevEx platforms, balancing immediate operational needs with strategic, future-state goals.
- Driving Cross-Team Alignment: Act as a critical liaison between infrastructure, security, and product development teams. Proactively drive cross-team alignment on architectural standards, tooling choices, and development workflows to ensure consistency and shared accountability for system health.
- Bottleneck Identification and Mitigation: Systematically identify engineering bottlenecks, friction points, and points of organizational toil within the SDLC. Implement targeted solutions—whether technical, process-based, or organizational—to mitigate these constraints and enhance overall engineering velocity.
- Planning and Execution: Collaborate with engineering leadership to transform the strategic roadmap into actionable, prioritized plans, securing cross-functional buy-in and resources for successful execution.
Qualifications and Education Requirements:
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 10+ years of relevant experience in software engineering, cloud architecture, and/or Site Reliability Engineering, with at least 3 years in a leadership or lead contributor role.
- Deep expertise of AWS, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA and S3.
- Strong proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation).
- Specialized experience in optimizing large-scale data platforms, specifically with Apache Spark. Proven ability to profile, troubleshoot, and tune Spark jobs for performance, cost, and reliability.
- 5+ years of experience with Kubernetes and containerization in general, including associated tools (kubectl, Helm, ArgoCD).
- Strong knowledge of AWS cost optimization.
- TCP/IP networking, including routing and AWS security groups.
- Excellent knowledge of CI/CD concepts and experience developing associated pipelines in CircleCI.
- Proficient in high-level scripting languages, including shell scripting, Python, and/or JavaScript.
- Experience with OTel and monitoring tools such as Splunk or DataDog. Experience with native AI observability tools is a plus.
- Experience with evaluating and rolling out GenAI tools for improving developer efficiency.
- Excellent communication, collaboration, and stakeholder management skills, with proven experience driving technical initiatives across multiple teams.
- Experience with researching and selecting new/modern developer toolsets and assisting teams in adopting them including vendor assessments, security assessments and procurement process.
- Experience in Ad-Tech or "BIG Data" processing organization is highly preferred
Target cash compensation range: $163,620 - $212,710 USD Annually
We are committed to providing competitive, market-informed compensation. The cash compensation above includes base salary, variable commission for employees in eligible roles, and annual bonus targets for eligible roles. In addition to cash compensation, all full time iSpotters are eligible to participate in iSpot's equity plan to receive stock options. Non-exempt roles will also be eligible for (pre-approved) overtime pay. Individual compensation packages are influenced by different factors unique to each candidate, including their skills, experience, qualifications and other job-related reasons.
For more information on total rewards package, go HERE
Hybrid & Flexible Workplace Policy
iSpot supports a hybrid and flexible workplace. Depending on location and work responsibilities, employees may be designated as full-time or part-time office-based or a fully remote employee. A hybrid work schedule indicates that you work in the office some days and work from home other days. The best hybrid workplaces allow for flexibility while also encouraging consistency.
Those local or living in surrounding areas to one of our offices (Bellevue, WA or New York, NY) will work a hybrid schedule, coming into their local office 1-3 days a week. While those in a role, not office-based and located further away from our offices, will work a fully remote schedule. If you have questions regarding exact details of our hybrid & flexible workplace policy, please let your recruiter know and they will discuss with you further.
#LI-Hybrid
If you don't feel you met every single requirement for the role, don't rule yourself out. Please apply anyway!
iSpot is an equal opportunity employer. All applicants will receive consideration for employment without regard to race, ethnicity, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. If you need assistance and/or a reasonable accommodation due to a disability during the application or the recruiting process, please contact our HR team.
California Residents applying for positions at iSpot can access our California Consumer Privacy Act here.
- ...Principal Site Reliability Engineer location- Washington, DC -Onsite Remote- No 6+ Months Job Summary At Amtrak, we're seeking a seasoned Principal Site Reliability Engineer with a strong focus on pipelines as code, CI/CD, and IaC. You...PrincipalRemote work
- ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...SuggestedTemporary work
$185k - $230k
As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver technical solutions within the highest levels of the federal government.We know that you can’t have great technology services...SuggestedFull timeLocal areaImmediate start$150k - $180k
...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business... ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or...SuggestedPermanent employmentFull timeWork at officeLocal areaRemote workWorldwide$165k - $230k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts....SuggestedPermanent employmentTemporary workImmediate startWeekend work$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine...Full timeWork experience placementImmediate start$207k - $284.9k
...on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is... ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers...Permanent employmentLocal areaWorldwideFlexible hoursDay shift$166k - $220k
...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &... ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril...Full timeWork experience placementImmediate start$112k - $179k
...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems...Contract workWorldwideShift work$125k - $185k
Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions... ...locate missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is...Remote work$125k - $185k
...lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package$210k - $230k
...Overview: GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing...Currently hiringRemote work- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment...Work experience placement
$174k - $239k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability...Local areaWorldwideFlexible hours$174k - $238k
...work. We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable,...Local areaWorldwideFlexible hours$82.3k - $228.8k
...inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer – Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud...Temporary workWork at officeRemote workWorldwide3 days per week$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...Work experience placementWork at office- ...The Role Join us in revolutionizing the lending landscape. SoFi is seeking enthusiastic Principal Software Engineers who are ready to lead the technical and strategic evolution of our financial services platform in support of our goals that put our members in control...PrincipalFull timeTemporary workWork experience placement
$153k - $185k
...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical...Permanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work$196k - $294k
...ABOUT THE TEAM ~ We are seeking a Principal Software Engineer to lead architectural and strategic initiatives for our ArsenalOS Forge platform. You'll be pivotal in evolving Forge's capabilities, driving integration and scalability, and maturing its architecture for...PrincipalFull timeWork experience placementLocal areaRelocation package$130k - $180k
...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...Temporary workWork at officeImmediate startRemote workFlexible hours- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...InternshipWork at officeLocal areaRemote workWorldwide
$100k - $110k
...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will...Permanent employmentRemote workFlexible hours- ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats...
- ...VirginiaHybridFull Time$180k - $200kAn early-stage AI startup is hiring a Principal Software Engineer to join its team. This is a hybrid opportunity based in... ...the operational challenges that emerge at scale - from reliability and observability to security, governance, and...Principal
- ...SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities... ...support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage,...Remote work
- ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise Cloud transformation...Temporary workImmediate start
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- principal network engineer Washington DC
- senior director engineering Washington DC
- civil engineer project manager Washington DC
- principal developer Washington DC
- principal engineer Washington DC
- director data engineering Washington DC
- principal infrastructure engineer Washington DC
- senior civil engineer project manager Washington DC
- director software engineering Washington DC
- general engineer Washington DC


