Cloud Reliability Engineer
Versant Media
Job Description
Job Description
Company Description
VERSANT (Nasdaq: VSNT) is an industry-changing media and entertainment business and home to trusted brands that shape culture, inform audiences, and build lasting connections. It operates across four core markets: political news and opinion, business news and personal finance, golf, and sports and genre entertainment. These markets are served through a powerful portfolio of iconic and innovative brands, including CNBC, MS NOW, USA Network, Golf Channel, Oxygen, E!, SYFY, and Versant's sports division USA Sports, along with complementary digital assets including Fandango, Rotten Tomatoes, GolfNow and GolfPass.
Job DescriptionThe Cloud Reliability Engineer is responsible for ensuring the availability, performance, scalability, and operational excellence of VERSANT’s cloud platforms and services.
This role works closely with cloud engineering, application development, networking, security, and operations teams to build and maintain highly reliable systems across a large multi-account AWS environment. The engineer will leverage automation, observability, and reliability engineering practices to improve platform resilience, reduce operational risk, and enhance the customer experience.
As a leading media company, VERSANT operates digital products, streaming platforms, content delivery systems, and media workflows that demand high levels of uptime and performance. The Cloud Reliability Engineer will help ensure these services remain resilient, scalable, and operationally mature.
The ideal candidate has strong experience with AWS, monitoring and observability platforms, incident management, automation, infrastructure as code, and operational best practices. Experience with AWS Organizations, Control Tower, Identity Center, Terraform, and modern cloud operations tooling is highly desirable.
Responsibilities
Reliabiliy Engineering
Design, implement, and maintain reliability practices for cloud infrastructure and platform services.
Define and monitor service-level objectives (SLOs), service-level indicators (SLIs), and operational metrics.
Identify reliability risks and implement solutions that improve availability, scalability, and resilience.
Drive continuous improvement initiatives focused on operational excellence and system stability.
Monitoring, Observability & Performance
Design and maintain monitoring, logging, alerting, and observability solutions across AWS environments.
Develop dashboards and reporting that provide visibility into platform health and performance.
Analyze system behavior, identify bottlenecks, and implement performance improvements.
Establish proactive monitoring practices that detect issues before they impact customers.
Incident Response & Operational Excellence
Participate in incident response, troubleshooting, and root cause analysis activities.
Lead post-incident reviews and identify corrective actions to prevent recurrence.
Improve operational processes, runbooks, and recovery procedures.
Support disaster recovery and business continuity initiatives.
AWS Platform Reliability
Support the reliability and operational health of large-scale AWS environments utilizing AWS Organizations, Control Tower, and Identity Center.
Partner with cloud engineering teams to improve platform architecture, resiliency, and operational consistency.
Assist in maintaining secure, scalable, and highly available cloud services.
Automation & Infrastructure as Code
Develop automation that reduces operational toil and improves system reliability.
Support infrastructure-as-code solutions using Terraform, CloudFormation, and related technologies.
Automate operational workflows, monitoring, remediation, and recovery activities.
Contribute to CI/CD pipelines and deployment automation initiatives.
Media & Digital Platform Reliability
Support the reliability of streaming platforms, content delivery systems, media workflows, APIs, and customer-facing applications.
Collaborate with engineering teams to improve application reliability and operational readiness.
Assist in capacity planning and scaling efforts for high-traffic events and media workloads.
Collaboration & Continuous Improvement
Partner with cloud, networking, security, and application teams to identify and address operational risks.
Promote reliability engineering best practices throughout the organization.
Contribute to documentation, standards, and operational procedures.
Evaluate emerging technologies and recommend improvements to platform reliability and observability.
Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
3–7 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or related roles.
Strong hands-on experience with AWS cloud services and enterprise-scale AWS environments.
Experience with:
Monitoring and observability platforms
Incident management and root cause analysis
Operational troubleshooting and performance tuning
AWS Organizations, Control Tower, and Identity Center
Experience with Infrastructure as Code:
Terraform
CloudFormation
Experience with CI/CD platforms and deployment automation.
Experience with scripting and automation using Python, PowerShell, Bash, or similar languages.
Strong understanding of AWS networking, resiliency, and cloud architecture concepts.
Experience with logging, metrics, tracing, and alerting technologies.
Strong troubleshooting, communication, and collaboration skills.
Additional Information
Location: New York City, NY or Englewood Cliffs, NJ - (Hybrid – 3 days onsite)
Employees based in our Englewood Cliffs office have access to a variety of convenient on-site services and amenities, including free employee parking, complimentary electric vehicle charging stations, an on-site fitness center with locker rooms and towel service, and dry-cleaning drop-off and pick-up.
To support an easier commute, shuttle service is also available to and from the Englewood Cliffs campus, with pickup locations across Manhattan’s East Side and West Side, Brooklyn, Hoboken, Jersey City, Newark, and Secaucus. Additional InformationAs part of our selection process, external candidates may be required to attend an in-person interview with a VERSANT Media employee at one of our locations prior to a hiring decision. VERSANT Media's policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.
If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to View email address on us.fitly.work.
VERSANT Media is committed to fair and equitable compensation practices. We include a good faith pay range for each position to comply with applicable state and local pay transparency laws and to promote equity across our organization. Actual compensation will be based on factors such as the candidate's skills, qualifications, experience, and location and may include additional forms of compensation and benefits such as health insurance, retirement plans, paid time off, etc.
VERSANT Media is not accepting unsolicited assistance from search firms for this employment opportunity. All resumes submitted by search firms to any employee at VERSANT via-email, the Internet, or in any form and/or method without a valid written Statement of Work in place for this position from VERSANT's Talent Acquisition team will be deemed the sole property of VERSANT. No fee will be paid in the event the candidate is hired by VERSANT as a result of the referral or through other means.
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
$124k - $186k
...where you can make an impact.With interesting opportunities in engineering, marketing, sales, supply chain, operations, HR, finance, and... ...have something special for you.POSITION SUMMARYThe Test and Reliability Engineer is responsible for ensuring we deliver product...SuggestedFull timeWork experience placementLocal area$30.7 - $46.05 per hour
...Equipment) required (safety glasses, gowning, gloves, lab coat, ear plugs etc.) Job Description Job Description The Engineer I, Reliability Engineering, at Thermo Fisher Scientific will play a crucial role in supporting equipment reliability, maintenance strategy...SuggestedHourly payTemporary workInternshipWork at office$99k - $132k
...Description Job Description Job Title: Reliability Engineer Location: Bayonne, NJ Department: Engineering Bayonne Experience: 5-7 years FLSA Status: Exempt Safety Sensitive: No We are a global leader in food & beverage ingredients. Pioneers at heart...SuggestedWork at office$123.8k - $175k
...customers to make the world healthier, cleaner, and safer through reliable manufacturing operations. You'll work with cross-functional... ...'s Degree plus 8 years of experience in maintenance or engineering in pharmaceutical/biotech manufacturing • Preferred Fields of...SuggestedTemporary workWork at office- ...Job Title: Reliability Test Engineer Location : Hoboken, NJ Division: Technology Department: Technology About Us Quantum Computing Inc. (QCi) (Nasdaq: QUBT) is an innovative, integrated photonics company that provides accessible and affordable quantum machines...Work at officeLocal areaRemote work
- ...world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office... ...with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and...Work at office
$90k - $120k
As a Performance II-Epic, your role is to provide reliability engineering services through observability and performance engineering techniques.... ...incidents. This role requires a strong background in scripting, cloud platforms, and a passion for optimizing operational...Full timePart timeWork experience placementRemote workFlexible hours$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online...Full timeTemporary workLocal areaRemote workWeekend work- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the AIML Platform team, you hold a leadership role in your team, demonstrate strong...
- ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology and Enterprise Technology Team, you draw upon your advanced knowledge to...Local area
- ...game-changing, high-quality solutions.As a Senior Lead Site Reliability Engineer at JPMorganChase within the Core Engineering Solutions team... ...an application or platform.Hands-on deep expertise in public cloud and modernization journey.Advanced knowledge and experience...
$117.6k - $160.73k
Job DescriptionAssociate Director, Cloud Computing & Systems EngineeringDivision of Information... ...Associate Director, Computing & Systems Engineering is hands on role that provides senior... ...automation, monitoring, security, reliability, and operational metrics· Enable cloud-native...Full timeFor contractors$50 - $65 per hour
...is seeking an experienced AI Platform Ops Engineer in Englewood, CO.Summary:The AI Platform... ...for someone with a strong background in cloud infrastructure, platform engineering, and... ...be responsible for maintaining a secure, reliable, and highly automated platform while partnering...$120k - $175k
...We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting... ...understanding of tooling and application development in these areas: Cloud computing such as AWS, Azure, and/or GCP. Infrastructure as...Full timeRemote workWork visaFlexible hours- ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering...Work experience placementRemote work
$113.3k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...Work at officeRemote workWorldwideFlexible hours- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...Local area
$119k - $170k
...founded in 2007 with a mission to make the cloud a safe place to do business and a more... ...make your next move with Zscaler. Our Engineering team built the world’s largest cloud security... ...looking for an experienced Staff Site Reliability Engineer (Federal) to join our...Full timeWork at officeLocal areaWorldwideNight shift- ...this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining... ...yourself among the top echelon in site reliability. As an Associate Site Reliability... ...payments, cybersecurity, machine learning, and cloud development. Our $9.5B annual investment...Worldwide
$130k - $180k
...at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy... ...experienced and innovative leaders and engineers in the field. Where we work Headquartered... .... The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team...Temporary workWork at officeImmediate startRemote workFlexible hours- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying... ...with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize...
- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the enterprise technology...
- ...Site Reliability Engineer As a Site Reliability Engineer, your role is to provide reliability engineering services through observability... ...incidents. This role requires a strong background in scripting, cloud platforms, and a passion for optimizing operational...Work experience placement
- ...develop game-changing, high-quality solutions. As a Lead Site Reliability Engineer at JPMorganChase within the Corporate sector, Enterprise... ...deploying, and supporting highly available services in a public cloud environment (AWS, Azure, or GCP); familiarity with cloud-...Work at office
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will... ...engineering craftsmanship and advanced proficiency across cloud platform engineering, observability, and performance and reliability...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...Azure DevOps Engineer Location: remote (nearby NY /NJ) Duration: 1+ year Qualification: Microsoft Azure, Azure DevOps... ...teams to implement CI/CD automation on-prem and in Azure cloud Implementation and troubleshooting of continuous build and deployment...Remote work
- ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing... ...Qualifications: ~4+ years of experience in cloud infrastructure engineering, platform engineering, or...Work at officeShift workDay shift
- ...Information Technology group delivers secure, reliable technology solutions that enable DTCC to... ...will have in this role:As a Lead DevOps Engineer within DTCC’s Technology Research &... ...CloudFormation for AWS)Strong knowledge of cloud platforms such as Azure, GCP or AWSDevOps...Remote workFlexible hours
- ...Fandango is looking for a SENIOR PLATFORM ENGINEER to build our next big thing in platform... ...diverse domains, offering expertise in AWS cloud infrastructure and resources, containers... ...automation solutions that enable rapid, reliable, and secure deployments Design and implement...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cloud Reliability Engineer. Be the first to apply!

