Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

EXIGER

Who We Are:

Exiger is the AI partner for supply chain and procurement automation. Its centralized 1EXIGER.AI platform allows organizations to manage their entire operating network, from parts to suppliers, regulators, and customers, through an autonomous agentic workforce that learns, adapts, and augments with every decision. The AI-native platform, combined with the largest supply chain knowledge asset, empowers 550+ global customers, including 150 Fortune 500 and 80+ government and defense industrial organizations, to act first. Exiger is FedRAMP® authorized and a 2x Leader in Gartner® Magic Quadrant™ for Supplier Risk Management.

Site Reliability Engineer

Location: U.S. (Hybrid)

This role requires U.S. citizenship and eligibility for a U.S. security clearance.

Role Summary:

Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI-powered platform equips the world's most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data-driven operational decisions.

From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high-stakes environments.

This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job.

You will use your expertise in coding, algorithms, complexity analysis, and large-scale distributed system design to solve the reliability challenges that are unique to operating a mission-critical AI platform in regulated and government environments.

SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

What You'll Do:
  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt.
  • Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight.
  • Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities.
  • Own the reliability of production services from design consulting and launch reviews through steady-state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self-service tooling.
  • Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault-injection testing, running game days that prove the platform degrades gracefully and recovers from failure.
  • Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on-call rotation.
  • Leverage AI-assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively.
What You Need:
  • Bachelor's or Master's degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function.
  • 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work.
  • A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics-driven decisions rather than relying on intuition or anecdote.
  • Experience with chaos engineering or fault-injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience.
  • Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems to root cause.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high-stakes environments.
  • Willingness to travel as needed to support customer engagements.
Nice to Have:
  • 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script.
  • Experience supporting ML or data platforms in production.
  • Familiarity with data warehouses such as Snowflake, Redshift and/or Apache Iceberg.
  • Experience operating in FedRAMP or other regulated or government environments.
Why You'll Love Working at Exiger:
  • High-performance culture rooted in accountability, collaboration, and a shared commitment to excellence.
  • Discretionary Time Off for all employees, with no maximum limits on time off
  • Industry leading health, vision, and dental benefits
  • Competitive compensation package
  • 16 weeks of fully paid parental leave
  • Flexible, hybrid approach to working from home and in the office where applicable
  • Focus on wellness and employee health through stipends and dedicated wellness programming
  • Purposeful career development programs with reimbursement provided for educational certifications

#LI-hybrid

Exiger is named a Leader in the Gartner® Magic Quadrant™ for Supplier Risk Management, twice selected as one of Fast Company's 'Brands That Matter,' and recipient of the Third Party Risk Association's Innovator Award, Exiger's technology has been recognized by leading analyst evaluations and 50+ awards. Learn more at Exiger.com and follow Exiger on LinkedIn .

At Exiger, our values define how we work and why we lead. We are mission-inspired, imagination-driven, trust-anchored, and compassion-focused-committed to building technology that makes the world safer, more transparent, and more resilient.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.

Exiger's hybrid work policy is periodically reviewed and adjusted to align with evolving business needs.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Jersey City, NJ vacancy
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex... 
    Suggested
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer... 
    Suggested
    Local area

    Hackajob

    Jersey City, NJ
    4 days ago
  • $137.75k - $185k

     ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Consumer and Investment Banking team, you will solve complex and broad business problems... 
    Suggested
    Local area

    JPMorgan Chase Bank, N.A.

    Jersey City, NJ
    4 days ago
  •  ...The selected colleague will work at an MUFG office or client sites four days per week and work remotely one day. A member of...  ...MUFG is seeking a highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and reliable web application environment... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work

    MUFG

    Jersey City, NJ
    2 days ago
  •  ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new opportunities... 
    Suggested

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Cloud Foundational Services team, you hold a leadership role in your team, demonstrate... 

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial Investment Banking team of Fraud Prevention, you will solve complex and broad business... 

    JP Morgan Chase

    Jersey City, NJ
    1 day ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Audit Technology, Production Management team, you hold a leadership role in your... 

    JP Morgan Chase

    Jersey City, NJ
    2 days ago
  •  ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise... 
    Temporary work

    Long Finch Technologies

    Jersey City, NJ
    a month ago
  •  ...Job Description Job Description BCforward is currently seeking a highly motivated SRE Software Engineer. Job Title: SRE Software Engineer Location: Jersey City, NJ Duration: Temp - 12 months   Job Description We are seeking a  Software Engineer-Other... 
    Temporary work

    BCForward

    Jersey City, NJ
    a month ago
  •  ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing...  ...improvement. Position Summary: The Senior Site Reliability Engineer acts as an advanced senior... 
    Work at office
    Shift work
    Day shift

    Bank of America Corporation

    Jersey City, NJ
    23 days ago
  •  ...appropriate Collaborates with other software engineers and teams to design and implement...  ..., test, and implement availability, reliability, scalability, and solutions in their applications...  ...customers Supports the adoption of site reliability engineering best practices... 

    3B Staffing LLC

    Jersey City, NJ
    1 day ago
  • As a Lead Site Reliability Engineer at JPMorgan Chase within the Public Cloud team, you will blend hands-on engineering with program leadership to promote platform stability, ensure consistent execution across SRE teams, and partner closely with Engineering and Product... 

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  • $42.5k - $70.5k

     ...formal training or certification in software engineering concepts, along with 3+ years of applied...  ...in SRE best practices, including reliability, scalability, performance, security, enterprise...  ...system architecture. We champion a site reliability culture by defining,... 
    Full time

    J.P. Morgan

    Jersey City, NJ
    9 days ago
  • $61k - $101k

     ...formal training or certification in software engineering concepts plus 5+ years of applied...  ...ability to independently deliver well-scoped reliability work and escalate when appropriate....  ...exceptional professionals for this Senior Lead Site Reliability Engineer opportunity within... 
    Full time
    Worldwide

    J.P. Morgan

    Jersey City, NJ
    2 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the AI Machine Learning and Data platform team, you hold a leadership role in... 

    JPMorgan Chase & Co.

    Jersey City, NJ
    6 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Chief Data & Analytics... 
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...we serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted...  ...scalability, and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    4 days ago
  •  ...Job Title: Site Reliability Engineer (SRE) Location: Wood Ridge, NJ Duration: 6 Months Position Overview We are seeking a highly skilled Site Reliability Engineer (SRE) to join our digital engineering and operations team. The ideal candidate... 
    Permanent employment

    PROLIM Corporation

    Wood Ridge, NJ
    3 days ago
  • $61k - $101k

     ...Requirements: We need 5+ years of experience in SRE, production engineering, platform reliability, or infrastructure operations at enterprise scale. We...  ...Chase brands. Our Public Cloud team is seeking a Lead Site Reliability Engineer who will combine hands-on... 
    Full time

    J.P. Morgan

    Jersey City, NJ
    6 days ago
  • $60 - $65 per hour

     ...SRE Engineer (W2) Jersey City, NJ (Onsite) 6 Months Contract to Hire Job Description: Proficient in application development skills for more than one technology as well as multiple design techniques. Working proficiency in development toolset to design... 
    Full time
    Contract work
    Work experience placement

    Pinnacle Group

    Jersey City, NJ
    4 days ago
  •  ...overworked staff, and outdated infrastructure have compromised reliable access to essential medications. We are addressing these...  ...on — and you set the standard for how infrastructure is built. Engineers write application code; you make sure it deploys reliably,... 
    Full time

    robotrx.ai

    Newark, NJ
    10 days ago
  • $60 - $67 per hour

    Site Reliability EngineerColumbus, OH - onsite from day oneContract to HireOnsite Interview 60% SRE & 40% EngineeringTop skills: Public Cloud...  ...Required:Formal training or certification in software engineering /Site Reliability Engineering concepts and 6+ years of applied... 
    Full time

    Pinnacle Group

    Jersey City, NJ
    1 day ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 
    Work experience placement

    AmNet Services

    Jersey City, NJ
    4 days ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 

    AmNet Services

    Jersey City, NJ
    4 days ago
  •  ...Release Train Engineer Encourages team and subsystem level Continuous Integration and Testing and Communities of Practice around SAFe, Agile, and Lean system engineering This will be at the request of assigned TOM or Program Manager at a sub task level Drive... 

    Zortech Solutions

    Jersey City, NJ
    1 day ago
  • $119k - $170k

     ...the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work
    3 days per week

    Zscaler

    Short Hills, NJ
    3 days ago
  • Company DescriptionSonsoft , Inc. is a USA based corporation duly organized under the laws of the Common wealth of Georgia. Sonsoft Inc. is growing at a steady pace specializing in the fields of Software Development, Software Consultancy and Information Technology Enabled...

    Sonsoft

    Jersey City, NJ
    2 days ago
  • $325k - $375k

     ...financial firm in Jersey City, NJ, is seeking a Principal Software Engineer (Algorithmic Trading System, Java) to play a key role in...  ...electronic trading using technologies such as Aeron, ensuring reliable, ultra-fast data transmission* Cutting-Edge Technology: Design... 

    KForce

    Jersey City, NJ
    2 days ago
  •  ...exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Sr Lead Software Engineer at...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    JP Morgan Chase

    Jersey City, NJ
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!