Site Reliability Engineer
$100k - $110kGrabJobs
Opportunity Overview: This is a remote-first role that may require travel to Boston, MA for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will bridge the gap between AWS cloud infrastructure, MERN stack applications, and large-scale data workflows. You will spend roughly 60% of your time on live incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning, and 40% on engineering automated solutions to eliminate operational toil. What you’ll do: Production Operations: Maintain the continuous uptime, scalability, and security of our AWS-hosted MERN applications and backend data architectures. Serverless Execution: Manage, optimize, and troubleshoot event-driven architectures running on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts. Data Pipeline Execution: Monitor scheduled PySpark data workflows, execute standard operating procedures (SOPs) for large-scale data ingestion, and rapidly triage, rerun, or patch failed data processing jobs. Incident Management: Participate in a collaborative on-call rotation to rapidly triage, debug, and mitigate live application outages and data flow bottlenecks. Healthcare Compliance: Maintain strict HIPAA, SOC2, and HITRUST compliance profiles across all runtime environments, storage systems, and data pipelines handling Protected Health Information (PHI). Toil Elimination: Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark pipeline recovery steps. Observability Engineering: Build specialized dashboards and alerts to monitor Node.js event loops, PySpark job execution stages, driver/worker memory leaks, and data pipeline throughput anomalies. Post-Mortem Culture: Lead blameless post-mortems for operational and data processing failures, translating system crashes into permanent structural fixes. What you’ll need: SaaS Platform Experience: Minimum of 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale. AWS Cloud Engineering: Deep expertise operating AWS core services, specifically AWS Lambda, Amazon ECS/EKS, Amazon EMR or AWS Glue (for Spark), EC2, VPC networking, IAM permissions, and CloudWatch. Automation & Data Languages: Professional competency in writing, debugging, and maintaining automation scripts and data tools using Python (including PySpark APIs) and Node.js. Data Operations: Experience managing and troubleshooting distributed data orchestration pipelines, ETL tools, message queues (e.g., AWS SQS/SNS, RabbitMQ), or stream processing frameworks. MERN Stack Operations: Deep understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering. Database Administration: Practical experience managing, sharding, indexing, and optimizing production-grade MySQL DB & Athena (RDS or self-hosted). Infrastructure as Code: Proven ability to deploy and maintain immutable infrastructure utilizing Terraform or OpenTofu. Healthcare Experience: Minimum 1 year working within HIPAA-regulated environments. Direct experience securing data-at-rest and data-in-transit containing sensitive patient records is preferred. Education & Experience: Minimum of 4 years of software/systems experience, with at least 1-2 years focused on live cloud operations and distributed data workflow management is preferred. Crisis Management: Calm under pressure with a methodical approach to identifying and isolating PySpark driver OOM (Out of Memory) errors or data corruption during high-stress outages. Attention to detail and effective communications skills will be critical in working with clients and internal stakeholders is preferred. Pay & Perks: Fully remote opportunity with about 5% travel Medical, dental, vision, life, disability insurance, and Employee Assistance Program 401K retirement plan with company match; flexible spending and health savings account ️ Flex Time Off + company holidays Up to 14 weeks of paid parental leave Pet insurance The salary range for this position is $100,000 to $110,000 annually; as part of a total benefits package which includes health insurance, 401k and bonus. In accordance with state applicable laws, Cohere is required to provide a reasonable estimate of the compensation range for this role. Individual pay decisions are ultimately based on a number of factors, including but not limited to qualifications for the role, experience level, skillset, and internal alignment. This role is not eligible for hire in: CA Interview Process*: Connect with Talent Acquisition for a Preliminary Phone Screening Meet your Hiring Manager! Design Interview(s) Cross Functional Interview *Subject to change About Cohere Health: Cohere Health’s clinical intelligence platform and agentic AI-powered solutions connect health plans’ strategic goals and providers’ needs, optimizing the speed, cost, and quality of care. With an enterprise approach that streamlines payer-provider decision-making across the care continuum–including policy, prior authorization, payment accuracy, and more–the company improves collaboration and reduces burden, resulting in up to 8x ROI and 94% provider satisfaction. With the acquisition of ZignaAI, we’ve further enhanced our platform by launching our Payment Integrity Suite, anchored by Cohere Validate™, an AI-driven clinical and coding validation solution that operates in near real-time. By unifying pre-service authorization data with post-service claims validation, we’re creating a transparent healthcare ecosystem that reduces waste, improves payer-provider collaboration and patient outcomes, and ensures providers are paid promptly and accurately. Cohere Health’s innovations continue to receive industry wide recognition. We’ve been named to the 2025 Inc. 5000 list and in the Gartner® Hype Cycle™ for U.S. Healthcare Payers (2022-2025), and ranked as a Top 5 LinkedIn™ Startup for 2023 & 2024. Backed by leading investors such as Deerfield Management, Define Ventures, Flare Capital Partners, Longitude Capital, and Polaris Partners. The Coherenauts, as we call ourselves, who succeed here are empathetic teammates who are candid, kind, caring, and embody our core values and principles . We believe that diverse, inclusive teams make the most impactful work. Cohere is deeply invested in ensuring that we have a supportive, growth-oriented environment that works for everyone. We can’t wait to learn more about you and meet you at Cohere Health! Equal Opportunity Statement: Cohere Health is an Equal Opportunity Employer. We are committed to fostering an environment of mutual respect where equal employment opportunities are available to all. To us, it’s personal. #LI-Remote #BI-Remote
$145k - $175k
...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data...SuggestedFull timeRemote work- ...our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-... ...for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving...SuggestedFull timeRemote workWork from home
$210k - $220k
...secure and private by design, it’s popular with security, IT, engineering, finance, and other security-focused teams. At Tines, we're... ...we’re looking for others to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team responsible...SuggestedWork at officeRemote work$141.8k
...their best work, grow fast, and bring their full selves to the herd. Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...SuggestedTemporary workRemote work$130k - $180k
...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...SuggestedTemporary workWork at officeImmediate startRemote workFlexible hours- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...InternshipWork at officeLocal areaRemote workWorldwide
- ...meaningful products that make a real impact on children's education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of Epic's platform as we grow. You are an experienced engineer who...Remote work
$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting – Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Work at officeRemote work- ...This individual must be based within either Eastern or Pacific Time in the US. Opportunity Deepgram is seeking a Pre-Sales Solutions Engineer to join our Applied Engineering team. While you'll have the opportunity to contribute across the team's various responsibilities (...Work at officeRemote workHome officeFlexible hours
$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting – Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Work at officeRemote work$105k - $140k
...Inclusivity. Today, nearly 200 people around the globe work on Speechify in a 100% distributed setting. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...- ...potential of our platform, guiding them from discovery to realization. Working closely with Enterprise Account Executives, our Solutions Engineers craft compelling product demonstrations, design and execute impactful proof-of-concepts, and shape technical strategies that drive...Remote work
- ...resilient EVP/TSP services. Requirements Requires Master's degree or foreign education equivalent in Computer Science or Computer Engineering + 3 years' experience in a software development role. Alternatively, Bachelor's degree + 5 years' experience. This is a...Remote workFlexible hours
- ...good advertising to thrive. About The Team We believe that engineers solve problems end-to-end. From researching the problem to... ...gatherings in New York City (our original home) and quarterly on-sites in various locations across the world in order to maintain the...Work at officeRemote workWork from homeWorldwide
$185k - $225k
...all of the time. About your role We're hiring Senior Software Engineers to join our AI-native engineering team. You'll design, build,... ...prompt/tool orchestration, and model observability • Improve the reliability, scalability, observability, and operational quality of...Full timeWork at officeLocal areaRemote workWork from homeFlexible hours- ...proprietary multi-modal dataset for training and evaluation Architect a deterministic secondary perception system Mentor junior engineers about best practices You have ~ Bachelor’s or Master’s degree in Computer Science, Robotics, Deep Learning, or a related...
$180k - $210k
...make the world safer and more secure. The Full-Stack Product Engineering team is responsible for building a comprehensive set of features... ...are looking for a pragmatic full-stack engineer who can build reliable and scalable software that will be used by all TRM customers....Immediate startWorldwide$75k - $150k
...industry , committed to making a positive impact on its customers, employees, and communities. The Role Veevaislookingforan Automation Engineer who is passionate about quality and automation. You will be creating, maintaining, and improving automation frameworks/...Work at officeLocal areaRemote workWork from homeFlexible hours- ...stability — on a system that demands high concurrency, high transaction volumes, and exceptional reliability. You will work cross-functionally with product, cloud, and engineering teams to deliver high- performance solutions that improve patient outcomes and modernize...Remote workFlexible hoursShift work
$160k - $240k
...foundation for sports predictions at scale, and we're looking for the engineers to help us do it right. The role We are building the future of... ...underpin our exchange infrastructure — focusing on latency, reliability, and accuracy across the most critical paths in our platform....Full timeContract workRemote workHome officeFlexible hours$114.8k - $191.4k
...invigorated and unstoppable with us! The Role As a Sr. Software Engineer on the Platform Engineering team at Bloomerang, you build and... ...and evaluation infrastructure that turns frontier models into reliable, everyday engineering leverage in service of our mission to empower...Permanent employmentFull timeLocal areaRemote workRelocation packageFlexible hours$108.08k - $112k
Employer: Omeda Holdings, LLC Job Title: Software Developer Job Code: # req 21144.2.6 Job Location: Vernon Hills, IL and various unanticipated locations throughout the U.S. Job Type: Full Time Rate of Pay: $108,077 – 112,000 per year Job Duties...Full timeRemote work- ...About Aptible Large language models are transforming software engineering — developers can now go from idea to code faster than ever.... ...complexity downstream: deployment, observability, cost, security, reliability — all of that has to scale too, and LLMs don't yet solve the...Local areaShift work
$180k - $260k
...systems and SaaS tools. Traditionally, acting on data has required engineering time and bandwidth, and left most business users stuck with... ...streaming sources like webhooks and queues Scalability and Reliability: As part of our rapid growth, we’re always evaluating current...Remote work$95k - $120k
...Description Our partner is looking for a Software Engineer for aRemote role. This is a full-time opportunity for a Software Engineer... ...architecture, collaborating with the engineering team to improve system reliability and functionality. The platform supports compliance workflows...Full timeRemote work- ...military base, Ditto's peer-to-peer sync engine ensures devices stay connected and data stays... ...with Robotic Platforms: Lead the on-site software integration of our platform with... ..., aerial, and maritime systems, building reliable data bridges between our synchronization...Fixed term contractLocal areaImmediate startRemote workFlexible hours
- ...reimbursement journey. About the Role We're hiring a Senior Software Engineer to join Pivotal's Growth & Onboarding team. This is a primarily... ...need at the earliest stages of the customer lifecycle into reliable, well-designed backend systems and workflows. Contribute to...Remote workFlexible hours
$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting - Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Full timeWork at officeShift work- ...meaningful and connected. If you’re excited to change how the world sells, join us. For more information, visit The Role Orum's Engineering team plays a pivotal role in driving this vision forward. A remote first company since the beginning- we're tackling some of the...Remote workFlexible hours
$170k - $220k
...power the future of the finance industry, we would love to hear from you! Who you are An ideal candidate is a product-minded engineer who wants to understand the business, the product, and the best ways to deliver value to our customers. You are someone who loves...Remote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Riverside, CA
- junior website developer Riverside, CA
- site leader Riverside, CA
- historic site Riverside, CA
- construction site safety Riverside, CA
- official site Riverside, CA
- site services specialist Riverside, CA
- site safety Riverside, CA
- IT site lead Riverside, CA
- site reliability engineer remote

