Director, Site Reliability Engineering
$189.59k - $220kNBCUniversal Inc
Director, Site Reliability Engineering
NBCUniversal is one of the world’s leading media and entertainment companies. We create world‑class content, distribute across film, television, streaming, and bring to life through global theme parks and consumer products.
Job Description
- As a member of NBCUniversal’s Production Software Engineering team, responsible for leading and performing custom architectural design, implementation, monitoring, and maintenance for a portfolio of production application environments.
- Responsible for hands‑on configuration and support as well as managing the work of other architects and engineers.
- Work closely with our Principal Software Engineer on technical architecture and design based on customer product requirements, translating product requirements to technical designs and implementations.
- Collaborate with cross‑functional team members such as Scrum Leads, Software Engineers, QA Engineers, UX Designers, Product Managers, other Architects & Site Reliability Engineers (Contractors and/or Staff), and third‑party vendors.
- Effectively delegate responsibilities to team members, mentoring and providing them with repeatable processes, and verifying the quality of their work.
- Utilize metrics to measure accomplishments and monitors progress, ensuring milestones and projects are completed on‑time.
- Communicate progress and the impact of solutions in technical terms to technology partners and in business terms to business partners.
- Establish a reputation as the subject matter expert for every tech stack used in Production Software Engineering applications and how they all fit together while keeping current with new technologies, developing innovative technical ideas, and generating proposals.
- Work with product teams to learn business objectives, development teams to plan platform needs, QA to understand test strategy, and SRE on environments and deployments.
- Participate in Scrums, demos, and other Agile ceremonies and ensure accurate and timely status updates to the team.
- Serve as primary interface with the NBCU Cyber Security team for all security‑related initiatives, patching, remediations, etc.
- Hands‑on commissioning, configuration, administration, documentation, and support for all on‑prem & cloud (AWS) environments (Servers, Storage, Databases, Networking, Security, etc.).
- Technical impact analysis, implementation, and monitoring of all cyber, technology audit, enterprise engineering, & IT (Databases, Monitoring, etc.) activities related to Production Software Engineering applications and platforms.
- Create and manage CI/CD pipelines using tool likes Cloud Formation, Foreman, Jenkins, Nexus, Rundeck, Ansible, and Puppet.
- Lead implementation of monitoring and reporting framework using tools like Grafana, Influx, Graylog/Splunk, Selenium, New Relic, and Icinga.
- Recognize and identify potential technical impacts of enterprise change controls which could affect our applications and customers.
- Help improve performance, scalability, and reliability.
- Build and maintain distributed infrastructure and automation.
- Solve problems quickly and automates processes for the future.
- Direct management of other engineers and architects (Contractors and/or Staff). 24x7x365 availability for production outages, emergencies, and deployments.
- 100% telecommuting is permitted for this role.
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, or related field (or foreign degree equivalent), plus 10 years of experience as a Software Architect or in a related occupation.
- Hands‑on systems engineering experience on Linux/Unix platforms.
- Experience with technical leadership and people management.
- Experience with Continuous Delivery and SDLC practices.
- DevOps principles, experience with operational tools (Ansible, Puppet, Chef, Terraform) and best practices for infrastructure (on‑prem or cloud) and software deployment.
- Operational experience with large‑scale applications.
- Experience with NoSQL data stores (MarkLogic, MongoDB, Cassandra, DynamoDB, Couchbase, PostgreSQL, etc.).
- Experience with a broad range of enterprise technologies.
- Experience building real‑time, large‑scale, low‑latency distributed systems.
- Experience with Agile tools like Jira, GitHub or similar.
- Additional skills (8 years experience required):
- Experience using AWS Cloud in a production environment.
- Experience with AWS IAM, EC2, RDS, S3, Lambda, batch and step functions.
Benefits
This position is eligible for company‑sponsored benefits, including medical, dental and vision insurance, 401(k), paid leave, tuition reimbursement, and a variety of other discounts and perks. Learn more about the benefits offered by NBCUniversal by visiting the Benefits page of the Careers website.
Compensation
Salary range: $189,592 – $220,000 per year
Full‑time: 40 hours/week
Equal Opportunity Statement
NBCUniversal’s policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.
If you are a qualified individual with a disability or a disabled veteran, you have the right to request a reasonable accommodation if you are unable or limited in your ability to use or access nbcunicareers.com as a result of your disability. You can request reasonable accommodations by emailing View email address on click.appcast.io.
For LA County and City Residents Only: NBCUniversal will consider for employment qualified applicants with criminal histories, or arrest or conviction records, in a manner consistent with relevant legal requirements, including the City of Los Angeles' Fair Chance Initiative For Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, where applicable.
#J-18808-Ljbffr- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...SuggestedFull timeWork experience placementRemote work
$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...SuggestedFull timeTemporary workLocal areaRemote workWorldwide$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the...SuggestedWork at officeLocal areaImmediate start$207k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives. Mentor other engineers and contribute to the engineering community through documentation...SuggestedFull timeWork at office$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad...Shift work
$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as a Site Reliability...Full time$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement$207k - $300k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Master's degree in Computer Science or Engineering.Experience mentoring engineers and cultivating... ...consensus across cross-functional teams.Site Reliability Engineering (SRE) combines...$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people... ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,...Full timeTemporary workPart time- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package- ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect...
$190k - $260k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving...Full timeWork experience placementWork at officeLocal areaRemote workHome office$150k - $220k
...teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping...Local area$182k - $250.8k
...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
- ...Title: Sr. Site Reliability Engineer (SRE) Location: New York City, NY - LOCALS ONLY Work Arrangement: Hybrid, 3 days Duration: 6-Month Contract Experience Range: 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10-15...Contract workLocal area
$130k - $200k
...who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll...Shift work- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- Skip To Main ContentBack to SearchRemote in United States of America: New YorkSite Reliability Engineering& 4 othersLooking for something else?Find a vacancy that works for you. Send us your CV to receive a personalized offer.Find me a jobWe are seeking a specialized Observability...
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling...Flexible hours
- ...SRE Engineer Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well) This is an SRE role supporting...
- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
- ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare...Relocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!
- senior chief engineer New York, NY
- associate director engineering New York, NY
- general engineer New York, NY
- project engineer assistant project manager New York, NY
- chief design engineer New York, NY
- principal infrastructure engineer New York, NY
- principal cloud engineer New York, NY
- assistant chief engineer New York, NY
- chief engineer New York, NY
- principal developer New York, NY

