Site Reliability Engineer
$145k - $175kGrabJobs
Full-time Description At Commence, we’re the start of a new age of data-centric transformation, elevating health outcomes and powering better, more efficient process to program and patient health. We combine quality data-driven solutions that fuel answers, technology that advances performance, and clinical expertise that builds trust to create a more efficient path to quality care. With human-centered, healthcare-relevant, and value-based solutions, we create new possibilities with data. We provide proof beyond the concept and performance beyond the scope with a focus on efficiencies that transform the lives of those we serve. With a culture driven by purpose, straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data platform. You will bridge the gap between engineering and operations—embedding reliability as a first-class concern from architecture through deployment. This role is built for someone who thrives when systems are under pressure and who treats an outage as a problem to be engineered away permanently, not just survived. Design, implement, and own observability infrastructure including metrics, logging, tracing, and alerting across distributed systems. Define and enforce SLOs, SLIs, and error budgets in partnership with product and engineering teams. Lead incident response: triage, coordinate remediation, conduct blameless post-mortems, and drive systemic fixes. Build and maintain CI/CD pipelines that support rapid, safe delivery of changes to production. Collaborate with engineering teams on infrastructure changes; able to read, modify, and contribute to existing infrastructure-as-code (Terraform or CloudFormation). Design and operate highly available, fault-tolerant systems—including auto-scaling, failover, and disaster recovery strategies. Reduce operational toil through automation; eliminate manual processes before they become habits. Collaborate with software engineers to establish reliability-first design patterns and review architectures for operational risk. Manage Kubernetes or container orchestration environments at scale. Ensure systems meet compliance and security requirements, particularly those applicable to healthcare data (HIPAA, SOC 2). Provide technical mentorship and guidance to engineers across the organization on reliability practices. Participate in on-call rotation with a commitment to continuously reducing the need for it. Qualifications 7+ years of experience in SRE, platform engineering, or DevOps roles. Exceptional problem-solving under pressure—demonstrated track record of diagnosing complex, high-stakes system failures and building durable solutions. Deep hands-on experience with AWS services including EC2, EKS/ECS, Lambda, RDS, S3, CloudWatch, and related tooling. Familiarity with infrastructure-as-code (Terraform or CloudFormation)—able to contribute to existing configurations. Experience designing and operating distributed systems with strict availability and latency requirements. Proficiency in at least one scripting or systems language (Python, Go, Bash, or similar) for automation and tooling. Experience with container orchestration (Kubernetes, ECS) in production environments. Expertise in observability tooling (OpenSearch, Prometheus/Grafana, or equivalent). Hands-on experience with CI/CD platforms (GitHub Actions, Jenkins, CircleCI, or similar). Proven ability to define and operationalize SLOs and error budgets. Experience with relational and NoSQL databases—performance tuning, replication, and backup strategies. Strong working knowledge of networking fundamentals: DNS, load balancing, VPCs, TLS. Excellent communication skills—able to translate technical risk into business impact for non-engineering stakeholders. Additional Requirements AWS Certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator). Experience in healthcare technology or other regulated industries (HIPAA, SOC 2, FedRAMP). Familiarity with chaos engineering practices and tooling. Experience with data pipeline reliability (ETL/ELT workflows, streaming systems). Exposure to AI/ML infrastructure and the reliability challenges unique to model serving. Familiarity with additional cloud platforms (Azure, Google Cloud). Contributions to open-source reliability or infrastructure tooling. *Commence' headquarters are in Virginia Beach, VA, however we are open to remote candidates in the following states: AZ, AR, CO, DE, FL, GA, IL, IN, KS, KY, MA, MD, MI, MS, MO, MT, NC, NE, NV, NY, OH, OK, PA, SC, TN, TX, VA, DC, WI, and WV* Work Environment/Physical Demands The work environment and physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. This is a remote position. While performing the duties of this job, the employee regularly works in a climate-controlled environment. Candidates must be able to sit, read, work on a computer, and watch a computer screen for extended periods of time. Occasionally required to stand, walk, use hands and fingers, kneel or crouch. Commence is an equal employment opportunity employer. All personnel processes are merit-based and applied without discrimination on the basis of race, color, religion, sex, sexual orientation, gender identity, marital status, age, disability, national or ethnic origin, military and veteran status or any other characteristic protected by applicable law. Commence.AI is committed to providing equal employment opportunities to all applicants, including individuals with disabilities. If you require a reasonable accommodation to participate in the application process due to a disability, please contact Human Resources at View phone number on click.appcast.io or View email address on click.appcast.io. Please note that unless you are requesting an accommodation, all applications must be submitted through our online application system. Salary Description $145,000-$175,000
$145k - $175k
...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data...SuggestedFull timeRemote work- ...and responsible for ensuring the availability, scalability, and reliability of systems and applications.What will be your responsibilities... ...using tools like Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-to-date with industry trends...SuggestedWork at officeRemote work
- The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams...SuggestedFull timeWork at officeLocal area
$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...SuggestedFull timeTemporary workRemote work- ...Job Description NBCUniversal Operations & Technology is looking for a Staff SRE, Playout Engineering to provide technical leadership to a team of Site Reliability Engineers. This team drives reliability, observability, and operational excellence for cloud-based...SuggestedFull timeWork at officeLocal area
$130k - $160k
...customers. If you thrive in a fast-paced, ideas-led environment, you’re in the right place.Why this job’s a big deal:Our team of engineers, at all levels, work with the business leaders in defining the product roadmap and come up with innovative solutions to grow the future...Full timeWork experience placementSummer workWork at officeRemote workFlexible hours$135k - $168k
...businesses, organizations, and government. Job Summary The AI Engineer is responsible for designing, developing, implementing, and... ...ownershipMonitor implemented AI solutions for performance, reliability, usage, cost, quality, and business value; recommend...Remote workFlexible hours$116.36k - $155.15k
...demonstrated knowledge and experience in system architecture and engineering disciplines. Specific technical knowledge of enterprise level... ...Amazon Web Services. Supports due diligence activities including site surveys, design, design review, bill of materials creation,...Full timeTemporary workRemote work1 day per week$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting - Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Work at officeRemote work- ...Participate in on-call for production infrastructure. Qualifications ~4+ years in DevOps, SRE, platform, or cloud infrastructure engineering. ~ Deep, hands-on AWS: You can reason about IAM trust policies, VPC networking, and cross-account access without a diagram....
- Our Benefits - Designed with You in Mind Comprehensive Health & Well-being Coverage From your very first day, you’ll have access to medical, dental, vision, and prescription drug coverage - ensuring you and your family stay healthy and protected. Generous Paid Time...Full timeImmediate start
- ...implement, and maintain web applications and infrastructure components.Duties and ResponsibilitiesUnder the direction of the Sr. VP, Engineering, the Back-End Web Developer is responsible for:Coding highly efficient and scalable softwareRefactoring and improving...
$150k - $200k
The Director, Platform Engineering, is the senior engineering leader responsible for the architecture, deliver, operations, reliability, and scalability of the company's enterprise Data Lakehouse platform. This role leads teams responsible for platform engineering and...Full time$18 - $25 per hour
...October 2026. If you’re interested in Solutions, Consulting and Engineering this is the place for you! Why Solutions Consulting &... ...skills Ability to collaborate well with others Must provide reliable transportation & housing Certain states and localities...Remote jobHourly paySummer workInternshipWork at officeImmediate startShift work$150k - $250k
WorldQuant develops and deploys systematic financial strategies across a broad range of asset classes and global markets. We seek to produce high-quality predictive signals (alphas) through our proprietary research platform to employ financial strategies focused on market...Casual workFlexible hours$40 per hour
...The Software Engineer Intern implements developer tools or product features on a rapid-release cycle. You will work in an agile development... ...Travel Requirements & Working Conditions Minimal travel Reliable internet access for any period of time working remotely and not...Remote jobSummer workInternshipSummer internshipWork at office- ...background aligns with future opportunities, we’ll reach out directly when formal applications become available. About Software Engineering Roles at Danaher Are you passionate about building real-world applications, writing clean code, and solving meaningful...Remote jobInternship
- We are seeking an experienced Salesforce Developer / Integration Lead with 7+ years of enterprise Salesforce delivery experience. The ideal candidate will have strong hands-on expertise in Salesforce development, Energy & Utilities Cloud, OmniStudio, integrations, and ...
$18 - $50 per hour
...0 GPA Perks: Employee discounts at our top customer sites Networking with our global leaders Mentorship from senior... ...the program Position Overview: As an Electronics Reliability Product Engineering Intern , you will support the development of Simcenter...Remote jobHourly payFull timeInternshipWork at officeLocal area- Job Description We have a terrific opportunity with our direct client near Norwalk, CT, for an AI Engineer. For this firm, a very well-established and successful company which is building internal AI capabilities, we're looking for an Associate Applied AI Engineer to...InternshipWork at office3 days per week
$175k - $200k
...architecting and building scalable, low-latency systems in Linux environments, collaborating closely with quantitative researchers & engineers, and driving next-generation simulation and execution platforms.ResponsibilitiesDesign, build, and enhance core infrastructure...Full timeCasual work$120k - $140k
...1B, F-1 OPT, and STEM OPT extension at this time. Your Role We are seeking a motivated and technically proficient Solutions Engineer to serve as a trusted advisor to customers throughout the sales process and beyond. In this role, you will collaborate with Sales...Remote jobFull timeH1b$18 - $50 per hour
...minimum 3.0 GPA Perks: Employee discounts at our top customer sites Networking with our global leaders Mentorship from senior... ...Industries Software is seeking a Digital Thread & Systems Engineering Intern to support the Capital PreSales organization. This role...Remote jobHourly payFull timeInternshipLocal area- Job Description Job Description The Boathouse is an award-winning creative agency with clients including Apple, Netflix, Verizon, SingleCare, Lovesac, Roomba and more. We are hiring an AI Software Developer to: Build applications and IT infrastructure for our...Hourly payFull time
- ...Job Description Job Description Applications Engineer Currently accepting applicants in NJ, NY, and CT. GameChange Solar is... ...specifications, reports Ability to perform work independently and reliably according to work schedules and assignments Capability for...Full time
- MS Teams Developer PlatformNote: MS Teams Developer Platform experience MandatoryMandatory Requirements3+ years of hands-on.NET Development experience1+ years of hands on development experience on MS Teams Developer PlatformHands-On Experience working on Team’s Client ...Work at office
$97.3k - $146k
...our financial products to serve our customers and their families when they need us most - now and in the future. As an AI Solutions Engineer, you will partner with business stakeholders to identify opportunities where Artificial Intelligence, intelligent agents,...Full timeWork at officeLocal areaRemote work- ...a primary focus on the Java business layer. Integrate with backend services over REST, WebSockets, and FIX to deliver responsive, reliable functionality to end users. Participate in the full software development lifecycle, including design, implementation, unit testing...Temporary workWork at officeRemote workFlexible hours
$180k - $220k
Software Engineer II, Market RiskStamford, CTThis role is with a financial services organization, joining a small team responsible for developing... ...office and two days remote.ScheduleFull time, hybrid (3 days on-site, 2 days remote).Compensation$180,000 to $220,000 per year.What...Work at officeRemote work- ...the intersection of financial services and enterprise software engineering, building and evolving mission-critical systems that serve... ...impactful role where your contributions will directly shape the reliability, scalability, and evolution of a core enterprise platform....Work at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Norwalk, CT
- site safety Norwalk, CT
- construction site safety Norwalk, CT
- official site Norwalk, CT
- site reliability engineer
- site reliability engineering manager
- junior site reliability engineer
- site reliability engineer sre
- site reliability engineer remote
- lead site reliability engineer



