Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

IPC, Inc.

DEPARTMENT: Product Engineering / Operational Readiness
REPORTING TO: Senior Manager, System Engineering and Integrated Product Support
OFFICE LOCATION: New York, NY
ROLE TYPE: Hybrid/Full-time

IPC is a global fintech company that puts people at the center of innovation. With a strong global footprint, we empower financial institutions and capital markets with advanced cloud-based trading communications and managed connectivity solutions.

Through our portfolio of communications and connectivity solutions, we focus on solving business challenges and adapting to regulatory changes in the fast-paced global financial markets. This enables our clients to maintain consistent market access, a strong competitive advantage, and enhanced operational efficiency.

Join a team that is dedicated to delivering groundbreaking products and making a significant impact on our clients' success.

Overview of the Team

The Network Services and OSS Engineering team is responsible for the reliability, monitoring, automation, and operational support of IPC's global technology platforms. The team supports mission-critical voice, network, infrastructure, and application services that power communication and connectivity solutions for the world's leading financial institutions.

Working across multiple regions and time zones, the team partners closely with Infrastructure Engineering, Product Engineering, Information Security, Network Operations, Service Delivery, and third-party technology providers to ensure highly available, scalable, and secure services.

Role Overview:

The Site Reliability Engineer II (SRE) is responsible for the architecture, development, implementation, administration, and operational support of IPC's enterprise monitoring, observability, event management, and automation platforms. The role ensures the reliability, scalability, performance, and operational visibility of IPC's global infrastructure, applications, cloud environments, and customer-facing services.

The position combines traditional Site Reliability Engineering responsibilities with software development, automation engineering, observability architecture, and ITSM integration expertise. The engineer serves as a technical leader responsible for monitoring transformation initiatives, event correlation, ServiceNow integrations, dashboard development, operational automation, and continuous platform modernization. The role partners closely with Product Engineering, Cloud Engineering, Infrastructure Engineering, Service Operations, Security, and Architecture teams to ensure operational readiness across all IPC products and platforms.

The ideal candidate possesses strong Linux administration expertise, experience with enterprise monitoring platforms, automation technologies, and a passion for improving reliability through engineering best practices.

Monitoring & Observability Engineering
  • Own and improve the monitoring and observability capabilities for assigned services and platforms, helping teams detect issues earlier, understand system health, and operate reliable customer-facing services.
  • Maintain and enhance monitoring across IBM Netcool, Splunk, Datadog, Grafana, Nagios, HPE NNMi, OpenTelemetry, and event-management platforms.
  • Apply established standards and patterns for metrics, logs, traces, synthetic checks, alerting, dashboards, and operational reporting across cloud, on-premises, and hybrid environments.
  • Partner with engineering teams to define practical monitoring and production-readiness requirements for new and changing services.
  • Use known reliability factors, service history, and precedent to resolve diverse observability issues, escalating or seeking review at key decision points.
Software Development & Automation
  • Build dependable tools and automation that reduce operational effort and make production services easier to support.
  • Develop and maintain Python automation, REST API integrations, reusable monitoring components, and event-processing workflows.
  • Create automated remediation and self-service capabilities for well-understood operational scenarios, with appropriate safeguards and review.
  • Integrate monitoring and operational checks into CI/CD pipelines and use Infrastructure as Code to support repeatable deployments.
ServiceNow & ITSM Integration
  • Own assigned integrations between observability platforms and ServiceNow, connecting alerts, incidents, configuration data, and operational workflows.
  • Implement and support Event Management, IT Operations Management, CMDB, service-mapping, ticket-generation, and incident-lifecycle integrations.
  • Advise colleagues on difficult integration and workflow issues and recommend solutions grounded in platform capabilities and operational precedent.
Platform Reliability & Service Ownership
  • Take end-to-end responsibility for the reliability and day-to-day health of assigned monitoring services, systems, or workstreams, working independently with review at key points.
  • Administer and improve Netcool ObjectServer, Event Gateway, probes, dashboards, correlation rules, suppression, deduplication, enrichment, and related platform components.
  • Improve availability, performance, capacity, security, and maintainability through tuning, upgrades, lifecycle management, patching, and technical debt reduction.
  • Define and track meaningful service indicators, objectives, alerts, and error budgets for assigned systems.
  • Play an active role in production support and incident response, restoring service quickly while turning operational learning into lasting reliability improvements.
  • Troubleshoot and resolve complex infrastructure, application, and monitoring incidents; coordinate recovery for assigned systems and contribute effectively during major incidents.
  • Lead or contribute to root-cause analysis, identify recurring failure patterns, and deliver corrective actions that reduce repeat incidents.
  • Maintain clear runbooks, support procedures, service documentation, and operational readiness evidence.
Cloud & Container Observability
  • Support AWS, Kubernetes, container, distributed-tracing, and OpenTelemetry observability for defined environments and services.
  • Help teams interpret service health, choose effective telemetry, and adopt reliable operating practices across hybrid platforms.
Collaboration & Technical Influence
  • Work closely with engineering, operations, security, service management, and senior partners to deliver pragmatic reliability outcomes.
  • Communicate recommendations clearly, using evidence and operational context to persuade stakeholders when priorities or approaches differ.
  • Provide practical guidance to colleagues on difficult technical matters and contribute to standards, roadmaps, evaluations, and proof-of-value initiatives.
  • Balance delivery, operational risk, security requirements, and long-term maintainability when planning improvements.
Experience & Professional Level

This Site Reliability Engineer II opportunity is suited to a fully qualified professional with approximately five years of relevant experience in reliability engineering, production operations, monitoring, observability, platform engineering, or automation. You will have room to own meaningful systems and improvements, work independently, influence experienced partners, and continue developing your technical depth with support and review at important points.

How You Will Make an Impact:

As an OSS Site Reliability Engineer II, you will take meaningful ownership of mission-critical systems and help make them more reliable, observable, and ready to scale. You will solve real operational challenges, improve the customer experience, and grow your expertise alongside experienced technical partners.

  • Own the day-to-day reliability and performance of mission-critical services, following issues through to durable improvements.
  • Build and improve modern observability through actionable metrics, logs, traces, dashboards, and alerts.
  • Automate repetitive operational work to reduce risk, increase consistency, and give the team more time for higher-value engineering.
  • Improve incident detection and recovery by sharpening alerts, supporting effective response, and applying lessons learned.
  • Strengthen operational readiness for new and changing services through practical reviews, testing, documentation, and launch support.
  • Collaborate with senior engineers and technical partners to prioritize reliability work, protect the customer experience, and expand your SRE skills.
  • Success means more dependable services, faster detection and recovery, smoother launches, and measurable growth in your ability to own and improve complex systems.
Essential Skills and Experience to be Successful in this Role:
  • Bachelor's degree in a relevant field, or equivalent practical experience.
  • Approximately 5 years of experience in site reliability, platform engineering, or production infrastructure, with independent ownership of assigned systems or workstreams.
  • Strong Linux skills and working knowledge of Windows environments.
  • Hands-on experience with monitoring, alerting, logging, and observability for production services.
  • Ability to automate operational work using Python or a similar language, APIs, and configuration-management tools.
  • Proven troubleshooting across applications, infrastructure, and networks in highly available environments.
  • Clear communication, familiarity with IT service management practices, and the ability to manage competing priorities.
Desired Skills and Experience:
  • Experience applying SRE practices, including SLIs, SLOs, and error budgets.
  • Experience with modern observability, APM, and distributed tracing tools (for example, Datadog, Dynatrace, New Relic, Splunk, Prometheus, Grafana, or OpenTelemetry).
  • Experience with cloud, containers, or hybrid environments, plus Infrastructure as Code and CI/CD practices.
  • Experience integrating ServiceNow or comparable ITSM platforms with monitoring, event management, or automated workflows.
  • Experience in fintech, telecom, or another mission-critical environment, with responsible use of AI-assisted engineering tools.
What's in It for You?

At IPC, your compensation is only part of the package. We are committed to investing in a range of programs and initiatives to improve the overall experience of our employees.

In addition to a collaborative, high-performing team environment, we're pleased to offer benefits including:

  • Medical, Dental and Vision
  • 100% Employer Paid Short/Long Term Disability, AD&D and Life Insurance Coverage
  • Limited & HealthCare Flexible Spending Accounts
  • Dependent Care Flexible Spending Account
  • Health Savings Accounts with Employer Contributions
  • Pet Insurance
  • Legal Insurance
  • Critical Illness, Hospital Indemnity and Accident Coverage
  • Financial Wellness Account Fiduciary and Training
  • Identity Theft Insurance
  • 401(k) and Roth plan with matching contributions
  • Flexible PTO, Sick Pay and Public Holidays
  • Additional Time off for Charity Work and Volunteering
  • Tuition Reimbursement
  • Certification Bonus Program
  • Access to “IPC University” our Internal E-Learning Platform
  • Structured Onboarding Training and Peer Mentor Support
  • Parental Leave Policy
  • Free Mental Health Wellness Programs, Tools, Coaching and Free therapy sessions
  • Employee Referral Scheme

Further information about your benefits will be provided during your onboarding process.

Additional Information:

At IPC, we believe that hybrid working creates an inclusive, flexible environment where employees can perform at their best, and teams can collaborate, innovate, and celebrate successes together. We spend around 60% of our time in the office and around 40% of our time working remotely. Some employees may be required to work from the office or client sites more than 60% of the time, if required by their role and/or client needs.

Your precise work schedule will be determined by you and your Line Manager before commencement of employment with IPC.

You can explore more about our culture, offerings and commitment on and

IPC's Work Culture:

The IPC work culture is one that fosters inclusion, prioritizes innovation, and maximizes potential.

We are a global ecosystem, full of diverse people that together made IPC what it is today.

Our strength as an organization is the sum of our different backgrounds, perspectives, skills and geographies; supported by an ironclad commitment to constructive dialogue and open-mindedness.

We live and breathe our commitment to innovation by embracing bold ideas, seizing new opportunities and striving for excellence. Our people have continued to deliver ground-breaking solutions to our clients for over 50 years.

#J-18808-Ljbffr
Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  •  ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also... 
    Suggested
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    3 days ago
  •  ...Ireland Full time Are you passionate about building reliable, scalable systems that power critical business solutions?...  ...can learn more about LexisNexis Risk at the Role As a Site Reliability Engineer (SRE), you will bridge software development and IT operations... 
    Suggested
    Full time
    Work from home

    RELX

    New York, NY
    1 day ago
  • $100k - $135k

     ...Model-Based Manufacturing, where context-aware production planning informs design in real-time—eliminating the disconnect between engineering and manufacturing. Our platform automates what can be automated and captures tribal knowledge where automation falls short,... 
    Suggested
    Work at office
    Flexible hours

    Dirac, Inc.

    New York, NY
    14 hours ago
  • $104k - $178k

    ## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in digital performance solutions, helping our advertiser and agency partners Verify the quality of their digital campaigns, Optimise... 
    Suggested
    Full time

    DoubleVerify

    New York, NY
    4 days ago
  •  ...Improve the reliability of mission critical solutions, applications, and platforms Software development for enterprises Continuous...  ...Shell Languages Powershell and Bash Windows and Linux Years of Experience: 5 Years of Software Engineering #J-18808-Ljbffr
    Suggested
    Work experience placement

    InterEx Group

    New York, NY
    5 days ago
  • $100k - $250k

     ...systems) can be yours. What you’ll do Improve observability, reliability and availability by defining and measuring key metrics....  ...function improvements. Educate, mentor and hold accountable the engineering team to improve the reliability of our systems and make... 
    Local area

    Kalshi

    New York, NY
    1 day ago
  • $171.6k - $223k

     ...development and learning. It allows us to scale easily, enabling our engineers to enhance attention on new features and capabilities. A key...  ...of members all over the world. Peloton is looking for a Site Reliability Engineer to create the tooling and services which simplify... 
    Temporary work
    Local area

    Peloton

    New York, NY
    2 days ago
  •  ...firm that has been operating at scale since 2014, with millions of global users and a reputation for rigorous engineering. The Role As Senior Site Reliability Engineer , you will own the infrastructure foundation that the entire engineering organization depends on.... 

    Harrison Clarke

    New York, NY
    14 hours ago
  • $120k - $150k

     ...allows each person to achieve personal success and add value to our teams and communities. We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY. About The Role Join our Platform Engineering team as a Site... 

    Piper Sandler

    New York, NY
    14 hours ago
  • $150k - $200k

     ...Defence and Government. They’re continuing their expansion of their New York (and Washington DC) engineering teams and looking for an Infrastructure Engineer / Site Reliability Engineer / Forward Deployed Infrastructure Engineer with a strong software engineering... 

    JLA Resourcing Ltd

    New York, NY
    14 hours ago
  •  ...Job Description Chariot’s engineering hire will be responsible for taking the Chariot platform to the next level. You will lead the evolution of our banking product, DAF product and technology strategy, along with building a growing team of software engineers. You will... 
    Work experience placement
    Work at office
    Work from home
    Monday to Friday
    Monday to Thursday

    Chariot

    New York, NY
    14 hours ago
  • $200k - $240k

     ...problems and help health systems deliver better care, we'd love to meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their hands directly into the systems that keep our healthcare... 
    Work at office
    3 days per week

    Socket

    New York, NY
    4 days ago
  •  ...Responsibilities Support and enhance the reliability, availability, and performance of...  ...service improvements. Work closely with engineering teams to implement, test, and deploy...  ...production environments, platform engineering, site reliability, DevOps, or infrastructure... 

    Selby Jennings

    New York, NY
    14 hours ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software applications that serve millions of customers and process billions of dollars in payments. We're looking for a seasoned... 
    Casual work
    Work at office
    Remote work

    ZIP

    New York, NY
    5 days ago
  •  ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    re-tool®

    New York, NY
    14 hours ago
  •  ...you’ll be building the future of financial infrastructure. As part of our global expansion, we’re looking for a hands‑on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next‑generation financial platforms. This is a high‑impact... 
    Remote work

    Longbridge Singapore

    New York, NY
    1 day ago
  • $120k - $160k

     ...and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production operations and software development as you improve, manage, and... 
    Work at office
    Local area

    The Voleon Group

    New York, NY
    14 hours ago
  •  ...Job Title: Site Reliability Engineer (SRE) Job Location: New York, NY Job Type: Contract Job description : # Owning infrastructure end-to-end designing, building, and scaling cloud and on[1]prem systems as the sole SRE # Architecting the migration... 
    Full time
    Contract work

    Staffingine LLC

    New York, NY
    2 days ago
  • $120k - $180k

     ...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises... 
    Permanent employment
    Full time
    Relocation package

    Raydar

    New York, NY
    a month ago
  •  ...Job Description Job Description Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well)... 

    IVID TEK INC

    New York, NY
    a month ago
  • $191k - $226k

     ...— and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads... 
    Remote work
    Work visa
    Flexible hours

    Garner Health

    New York, NY
    a month ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure the company’s sophisticated distributed... 
    Work experience placement

    United States Digital Space LLC

    New York, NY
    2 days ago
  •  ...right care, at the right time, in the right setting. Role Description You'll join Tennr's Infrastructure team as a Site Reliability Engineer, focused on the systems that keep us reliable, observable, and secure. This is a hands-on role with real room to grow: you... 
    Work at office

    Tennr

    New York, NY
    14 hours ago
  •  ...visible impact and put something genuinely rare on your CV, keep reading . About The Role We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'... 
    Remote work
    Worldwide
    Flexible hours

    Camunda

    New York, NY
    2 days ago
  •  ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE I) to join our Video Platform Engineering Team. As a Level 1 SRE, you will work closely with senior engineers to respond to incidents... 
    Temporary work

    Elite Technical

    New York, NY
    a month ago
  • $139k - $257.55k

     ...The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems Inc

    New York, NY
    5 days ago
  • $120k - $140k

     ...Site Reliability Engineer – AWS, Azure, IaC, Typescript, .NET – NY (office 3 days a week) - $120,000 - $140,000 Do you want to be the first technical hire in the US for a fast-growing leading FX hedging platform? Do you want to be the cornerstone for the team as it... 
    Full time
    Work at office
    Rotating shift
    3 days per week

    Stott and May

    New York, NY
    4 days ago
  • $135k - $160k

     ...About the Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, all while keeping operational toil to a minimum for the Fabric tech community. What You’ll Do... 

    Fabric Labs

    New York, NY
    14 hours ago
  • $250k - $300k

     ...skills and experience — talk with your recruiter to learn more. Base pay range $250,000.00/yr - $300,000.00/yr Senior Site Reliability Engineer – Trading Systems This isn’t a support role. It’s an engineering position where uptime, performance, and automation... 
    Full time
    Work at office
    Home office

    Quantitative Systems

    New York, NY
    14 hours ago
  • $115k - $125k

     ...headquartered in New York, with offices in Chicago, London, Singapore, Hong Kong and Tokyo. Purpose of the role Pico Site Reliability Engineering is a customer-facing group engaged with our customers, development and production teams to ensure customer success with... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Work from home
    Monday to Friday
    Flexible hours
    Shift work
    Weekend work
    Afternoon shift
    Early shift

    Pico

    New York, NY
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!