Principal Site Reliability Engineer
$200k - $250kDraftKings
At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together.
The Crown Is Yours
As a Principal Site Reliability Engie r , you'll shape the long-term strategy for the infrastructure behind one of the most demanding platforms in sports betting and gaming. You'll drive the architectural direction of our cloud and on-premise platforms, helping engineering teams build, deploy, and operate highly reliable systems at scale. Working across Platform Engineering and Site Reliability Engineering, you'll influence how we modernize our infrastructure, strengthen operational excellence, and prepare our platform for the next generation of growth.
What you'll do
- Define and execute the long-term strategy for our Kubernetes platform across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise environments, ensuring reliability, scalability, and operational consistency.
- Drive architectural decisions across critical infrastructure, including cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost optimization.
- Lead large-scale platform initiatives across multiple engineering teams, establishing technical direction, engineering standards, and measurable outcomes that improve platform reliability and developer experience.
- Establish and evolve reliability practices by defining service level objectives, service level indicators, and error budget frameworks that align platform performance with business priorities.
- Build automation-first infrastructure through Infrastructure as Code, GitOps workflows, self-healing systems, and internal platform tooling that improve engineering velocity and reduce operational overhead.
- Champion the responsible adoption of AI-powered engineering capabilities that improve operational efficiency, accelerate incident response, and enhance developer productivity.
- Lead critical platform incidents, drive post-incident improvements, and strengthen platform resilience through automation, capacity planning, and operational excellence.
- Mentor senior engineers, influence technical strategy across the organization, and elevate engineering excellence through architecture reviews, coaching, and technical leadership.
What you'll bring
- A Bachelor's Degree in Computer Science or a related technical field.
- At least 8 years of experience designing, operating, and scaling distributed cloud and on-premise infrastructure, including at least 3 years operating at the Staff, Principal, or equivalent technical leadership level.
- Proven experience leading large-scale infrastructure or platform initiatives that require cross-functional alignment and long-term technical ownership.
- Deep expertise with Kubernetes, including cluster architecture, networking, storage, security, operators, lifecycle management, and large-scale production operations.
- Extensive experience building and operating production infrastructure in AWS and Google Cloud Platform using Infrastructure as Code technologies such as Terraform, Pulumi, or similar tools.
- Strong software development experience in Go, Python, or both, with expertise in GitOps, continuous integration and continuous delivery, observability, distributed systems, Linux, and reliability engineering principles.
- Experience incorporating AI-powered tools into engineering workflows while applying sound judgment around reliability, security, and operational risk.
- Exceptional communication and leadership skills with a proven ability to mentor engineers, influence technical strategy, and drive engineering excellence. Experience working in regulated industries, hybrid cloud environments, contributing to open-source projects, or holding cloud certifications is preferred.
Join Our Team
We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role.
The US base salary range for this full-time position is 200,000.00 USD - 250,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
#J-18808-Ljbffr$169.3k - $304.7k
...building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...PrincipalWork experience placementWork at office- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...SuggestedFull timeFlexible hours
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...SuggestedTemporary workWork at officeLocal areaFlexible hours3 days per week$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...SuggestedWork experience placementWork at officeRemote work$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...Work experience placementWork at officeRemote workFlexible hours$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Remote work$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that... ...maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering...Local areaWorldwideFlexible hours- ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the... ...ideal candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based...Full time
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week$160k - $200k
...promQL Key Responsibilities: Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coach Perform incident...Temporary workWork at officeLocal areaFlexible hours3 days per week$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$140k - $210.9k
...States. The position will be primarily on-site with residency commutable to one of our... .../DevOps backgrounds or software engineering backgrounds (e.g., Java Python, Go) with... ...strong interest in operating and improving reliability of distributed production systems. Responsibilities...Full timeTemporary workPart timeWork at officeShift work$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...Work experience placementWork at office$130k - $150k
...systems and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud...Work at officeWork from home3 days per week$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work$90k - $110k
...SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocations: Boston/Waltham, MA | HybridGet To Know us:SS... ...out talented candidates for the position ofSite Reliability Engineer. This role is based out of one of our Boston-area offices (Boston...Ongoing contractFull timeCasual workWork at officeWorldwideFlexible hours- ...we empower creators to own their own destiny. Be Klaviyo’s senior IC for scale, you will report into a VP of Engineering and lead performance, reliability, multi‑region, and large‑tenant readiness. You’ll drive platform-wide architectural change, hunt bottlenecks and...PrincipalFull time
$266.2k - $425.9k
...Acceleration group empowers over 2,000 engineers to build, test, and deploy their code at... ...incident management and production reliability. This is not a team focused on maintaining... ...process.About the RoleWe're looking for a Principal Software Engineer to help define the...PrincipalLive outWork at officeRemote workShift work- ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment... .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools...PrincipalFull timeRemote workRelocation packageFlexible hours
$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work- ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed...
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Information Technology group delivers secure, reliable technology solutions that enable DTCC to... ...This RoleAs a Senior Application Support Engineer, you will help power DTCC's global... ...trade processing and settlement.Leveraging Site Reliability Engineering (SRE) principles,...Remote workFlexible hours
- ...Software Development Engineer We're creating a platform that will change the way organizations measure their software development efforts... ...teams can work and the tools they use Location: on-site in Boston We believe that it takes a diverse team to build the...Principal
$146.25k - $225k
...to solving complex problems and making a huge impact. We are expanding our team and recruiting for a skilled Senior/Staff Site Reliability Engineer focused on designing, building, and operating our on-prem/cloud environment.The OpportunityYou will advance the state of our...Local areaRemote workRelocation package$118.3k - $147.9k
...CMT is looking for a Senior Site Reliability Engineer, SecOps to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — understanding...Full timeTemporary workSummer workWork from homeWorldwideFlexible hours$160k - $225k
...Staff Site Reliability Engineer Cambridge, MA Manifold is the AI platform for life sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification and clinical development to market access and precision...$150k
...MassachusettsHybridDirect Hire$145k - $165kOur client is seeking a Principal DevOps Engineer to join their team in Downtown Boston. This is a full-time... ...and build automated cloud infrastructureImprove system reliability, scalability, and performanceLead technical direction and...PrincipalFull timeRemote work$105.79k - $141.05k
...delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations teams and... ..., and automation, and you'll coordinate across architecture, engineering, and systems development organizations to measurably improve...Temporary workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- senior chief engineer Boston, MA
- associate director engineering Boston, MA
- general engineer Boston, MA
- principal infrastructure engineer Boston, MA
- principal cloud engineer Boston, MA
- chief engineer Boston, MA
- principal developer Boston, MA
- senior principal engineer Boston, MA
- engineering director Boston, MA
- principal data engineer Boston, MA


