Site Reliability Engineer
Comcast
Make your mark at Comcast -- a Fortune 30 global media and technology company. From the connectivity and platforms we provide, to the content and experiences we create, we reach hundreds of millions of customers, viewers, and guests worldwide. Become part of our award-winning technology team that turns big ideas into cutting-edge products, platforms, and solutions that our customers love. We create space to innovate, and we recognize, reward, and invest in your ideas, while ensuring you can proudly bring your authentic self to the workplace. Join us. You’ll do the best work of your career right here at Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual due to the nature of their work. If a position is listed with both office locations and virtual offerings, Comcast may be willing to consider candidates who live greater than 100 miles from the office for the remote option.) Job Summary COMCAST Technology Solutions is a software technology company headquartered in Denver, Colorado, USA. We enable streaming services, TV stations, pay TV operators, content providers, broadband media sites, and mobile businesses to solve their unique media management and video publishing requirements. Our Cloud Video Platform (CVP), provided as a service, offers a diverse product catalogue. By leveraging Comcast CVP, our customers can securely manage their digital media, publish content to various IP devices, and effectively monetize their distribution directly to consumers. Our proven media management and publishing technology provides a versatile approach to meet each customer’s unique business requirements and scales fluidly to support their growth. Our customers include Deutsche Telekom, Viaplay, Fox, Disney, NBC, Paramount+, and many others. Our Site Reliability Engineering (SRE) team is at the heart of our mission to deliver seamless and robust services to our users. We’re a distributed team of engineers with diverse skillsets who thrive on solving complex challenges and driving innovation with a focus on improving observability and reducing toil. Job Description *This position is unable to provide work authorization sponsorship or immigration support now or in the future.* Core Responsibilities * Analyzes and forecasts system capacity requirements to ensure scalability and performance for high-profile events. * Participate in incident response efforts, conduct post-incident reviews, and implement corrective actions and improvements to monitoring.
- Develop and maintain monitors and alerts across all services.
- Optimize system performance through tuning and configuration adjustments.
- Develop and maintain disaster recovery plans and procedures to ensure
- Shows regular, consistent and punctual attendance.
- Other duties and responsibilities as assigned.
ABOUT YOU
Our people are the most important part of our business. We are fundamentally looking for forward-thinking, enthusiastic problem solvers. People who love a challenge, constantly evaluate and question, and, above all, love to ship a product that solves real problems. While these characteristics outweigh any specific technical skills, you should be able to demonstrate some of the following * An understanding of wider operational performance factors influenced by the underlying infrastructure workload, such as server platforms, databases and networking. * A strong drive to be a ‘detective’ and understand why things are working (or not working) as they should, in other words, a passion for detail and an investigative nature. * The ability to proactively diagnose problems using your holistic knowledge-set – and then get busy with coding a permanent fix, rewriting a process or working with third parties to ensure that lessons are learned, and problems never recur. * A vision of automation as an opportunity to overcome scale challenges, and a flexible approach to technologies. Must Have Skills:- Strong knowledge of cloud platforms (e.g., AWS, GCP, Azure).
- Proficiency in scripting languages (e.g., Python, Bash).
- Experience with infrastructure-as-code tools (e.g., Terraform, Ansible).
- Familiarity with containerization and orchestration tools (e.g., Docker,
- Experience with monitoring tools (e.g., Datadog, Splunk)
- Experience with Database performance monitoring and tuning (e.g., NoSQL, SQL)
- Experience with Kubernetes performance monitoring and tuning
- Excellent problem-solving skills and attention to detail.
- Effective communication and collaboration skills.
- CI/CD pipeline management
- Cloud Cost Optimization
- Automating deployment processes
- Front-end development (e.g., React)
- Medical & Dental
- 401(k) Savings Plan
- Generous paid time off
- Life Milestones - from adoption assistance, childcare resources, pet
- Comcast is an EOE/Veterans/Disabled/LGBT employer.
- Discount tickets for Universal Resorts, including theme park tickets and
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- site reliability engineer remote United States
- site safety United States
- site merchandiser United States
- after school site coordinator United States
- historic site United States
- IT site lead United States
- site leader United States
