Site Reliability Engineer
$130k - $160kiCapital
About the Role
The Site Reliability Engineering team at iCapital is fundamental to ensuring our platform delivers consistent, reliable service to our client base. As a Site Reliability Engineer, you'll work at the intersection of software engineering and operations, applying engineering principles to infrastructure challenges. You'll be responsible for designing and implementing systems that scale efficiently, architecting observability solutions that provide actionable insights, and building automation that enhances our platform's reliability. This role requires someone who thinks systematically about reliability, can translate business requirements into technical implementations, and thrives on making complex systems more robust.
Responsibilities:
- Define, implement, and iterate service level objectives (SLOs) and service level indicators (SLIs) that reflect customer and business expectations.
- Lead monitoring and alerting standardization through “monitors as code” (Terraform preferred), including quality gates such as severity, ownership, and runbook links.
- Develop observability standards across metrics, logs, and traces, including instrumentation and dependency mapping patterns (OpenTelemetry where applicable).
- Lead technical evaluations and PoCs for observability platforms and integrations; define success criteria and migration approach for adoption.
- Define and implement reliability and operability standards for Kubernetes-based services, including scaling patterns, resource constraints, rollout safety, and baseline dashboards and alerts as part of service onboarding.
- Drive automation to eliminate toil, improve repeatability, and accelerate recovery (incident workflows, runbooks, and remediation where appropriate).
- Serve as Incident Commander for high-severity incidents, lead postmortems, and drive systemic improvements through action items and measurable follow-through using established tooling workflows.
- Participate in on-call rotations with a focus on improving reliability, reducing alert noise, and increasing signal quality over time.
Qualifications:
- 7+ years in SRE or related roles, with evidence of technical seniority across multiple services and teams.
- Strong experience with AWS and container orchestration (Kubernetes) in production environments.
- Demonstrated experience defining SLOs/SLIs and using them to drive operational and engineering decisions.
- Proven ability to design and implement observability solutions that produce actionable insights while reducing alert fatigue and operational noise.
- Strong IaC skills (Terraform preferred) and the ability to build reusable automation and standards (monitoring as code, configuration patterns).
- Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems.
- Experience with at least two observability stacks (Prometheus/Grafana, New Relic, Splunk, CloudWatch, ELK, etc.) and driving standardization across them.
- Strong incident response skills, including leading retrospectives/postmortems and improving reliability through systematic follow-up.
- Strong debugging skills across distributed systems and production environments, including performance and reliability investigations.
- Clear written and verbal communication skills with the ability to influence engineering teams through standards, tooling, and practical guidance.
Benefits
The base salary range for this role is $130,000 to $160,000 depending on level. iCapital offers a compensation package which includes salary, equity for all full-time employees, and an annual performance bonus. Employees also receive a comprehensive benefits package that includes an employer matched retirement plan, generously subsidized healthcare with 100% employer paid dental, vision, telemedicine, and virtual mental health counseling, parental leave, and unlimited paid time off (PTO).
We believe the best ideas and innovation happen when we are together. Employees in this role will work in the office Monday-Thursday, with the flexibility to work remotely on Friday.
For additional information on iCapital, please visit Twitter: @icapitalnetwork | LinkedIn: | Awards Disclaimer:
iCapital is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, gender, sexual orientation, gender identity, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.
- ...thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage...SuggestedFull timeShift work
- A recruitment agency for technical graduates seeks candidates for a program linking recent graduates to leading global employers in technology. Applicants will undergo rigorous training and support while working in production support roles, gaining valuable experience ...Suggested
- ...Position Name: Site Reliability Engineer Location: Salt Lake, UT USA Full time position Detailed Job Description. Well versed in Application Monitoring tools (Splunk, Extrahop, App Dynamics, Prometheus & Grafana) Good understanding...SuggestedFull time
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...SuggestedWork experience placementWork at office$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform....Full timeShift work
$84.9k - $209.5k
...Job Description As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve...Temporary workFlexible hours$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaWorldwideFlexible hours$160.8k - $214.1k
...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield... .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the...Full timeTemporary workLocal areaRemote workFlexible hours- ...Lake City, Utah, United States of America Job Title: Operations Reliability Scientist Reporting Relationships Direct Supervisor: VP of... ...process and product technical guidance to program managers and engineers, including training/mentoring manufacturing leaders. Act as the...
$73.8k - $110.7k
...mission to expand access to high-quality, affordable education. Our engineering teams build the platforms, services, and tools that support... ...systems, working with cloud-native technologies, and building reliable software solutions that support learners and employees. What...Full timeInternshipFlexible hours$125k - $191.7k
...Job Description Hybrid: This role is categorized as hybrid/Remote Role: As a Senior Software Systems Engineer on the Software Validation team within the AV organization, you will play a critical role in leading the strategy and execution of validation efforts...Local areaRemote workWork from homeFlexible hours- ...Description SUMMARY Responsible for ensuring the reliability, efficiency, and continuous improvement of manufacturing... ...Requires collaboration with cross-functional teams, including engineering, maintenance, operations, and quality, to achieve optimal equipment...Full time
$136k - $184k
Platform Infrastructure Engineer IV or V, DOEHybrid (Office 3 days/wk - Onsite-Flex) within Portland, OR; Medford, OR; Renton, WA; Burlington... ...platform expertise in cloud, networking, automation, and reliability engineering—tackling complex technical challenges, influencing...Full timeWork at officeImmediate startWork from homeFlexible hours- ...opportunity. We're looking for a talented and motivated Software Engineer II to join our Web Platform Team and help build the... ...infrastructureMonitor platform health and continuously improve reliability, scalability, and developer productivityContribute to an Agile...Full timeFlexible hours
$103.71k - $138.28k
...demonstrated knowledge and experience in system architecture and engineering disciplines. Specific technical knowledge of enterprise level... ...Amazon Web Services. -Supports due diligence activities including site surveys, design, design review, bill of materials creation,...Temporary workRemote work$293.9k - $406.8k
...the TeamYou will join Cisco’s Identity Engineering Group, a foundational organization responsible... ...domain, with a strong emphasis on reliability, interoperability, and long-term scalability... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workLocal areaRemote workFlexible hours$183.8k - $263.6k
...orchestration, and secure service integration. You will work closely with engineers across control plane, data plane, and platform teams to deliver... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaRemote workFlexible hours- ...seeking a highly skilled Senior Backend-Focused Full Stack Software Engineer to design, build, and scale modern applications and distributed... ...engineers, and engineering leadership to deliver secure, reliable, and high-performing solutions. This is an opportunity to influence...Full timeFlexible hours
- What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build...Full timeTemporary workPart timeImmediate startRemote work
- ....00 - $178,500.00Job DescriptionAt WGU, technology powers opportunity. We're looking for a talented and motivated Mobile Software Engineer II to help build and enhance the next generation of WGU's student mobile experience.This role is ideal for an engineer who enjoys...Full timeFlexible hours
- Salt Lake City, UtahEngineering - Engineering /Full-time /On-siteFilevine is a Legal AI company delivering Legal Operating Intelligence for... ...).Willingness and ability to travel extensively to client sites nationwide.Bachelor's degree in Computer Science, Engineering,...Full timeTemporary work
- ...estimate of the current range is: Grade: Technical 408Pay Range: $118,900.00 - $178,500.00Job DescriptionWe're looking for a Salesforce Engineer to join the Student Lifecycle Services engineering team at WGU. You'll work on a large, complex Salesforce org that directly...Full timeWork at officeFlexible hours
- ...for companies who are owned or affiliated with The Church of Jesus Christ of Latter-day Saints.We are seeking an Appian Software Engineer to join our Platforms team. You'll build and enhance secure enterprise applications that support healthcare administration, business...
$130k - $170k
Company Founded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the digital... ..., Bedrock, and Lambda Architect multi-agent systems with reliable tool use, memory, and orchestration patterns that perform accurately...Full timeWork at officeWork from homeFlexible hours- ...technology solutions connecting the space, air, land, sea and cyber domains in the interest of national security. Job Title: Systems Engineer, Lead Job Code: 41201 Job Location: Salt Lake City, UT Job Schedule: 9/80 work- employees work 9 out of 14 days- totaling...Local area
$122.5k - $423.78k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelDirectorJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions for clients. They play a...Full timeTemporary workH1bRemote work- ...the interest of national security. Job Title: Lead Systems Engineer Job Code: 38278 Job Location: Salt Lake City, UT Job Schedule... ...Engineer to be responsible for evaluating and improving the Reliability of Department of Defense (DoD) communications systems (...Local area
- ...connecting the space, air, land, sea and cyber domains in the interest of national security. Job Title: Lead, Networking Systems Engineer Job Code: 38859 Job Location: Salt Lake City- UT Job Schedule: 9/80 Employees work 9 out of every 14 days- totaling 80 hours working...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site services specialist Salt Lake City, UT
- construction site safety Salt Lake City, UT
- site leader Salt Lake City, UT
- official site Salt Lake City, UT
- website content developer Salt Lake City, UT
- on site coordinator Salt Lake City, UT
- IT site lead Salt Lake City, UT
- site safety Salt Lake City, UT
- junior website developer Salt Lake City, UT
- on-site clinical research associate (traveling/remote) Salt Lake City, UT

