Site Reliability Engineer
$145k - $165kBolt Graphics
Bolt Graphics is a semiconductor startup based in Sunnyvale, CA building the fastest and most efficient graphics processors. We pride ourselves on our first principles approach to solving problems. We are energized by our mission to reduce the barrier of entry for content creation and consumption. Our goal is to enable everyone to easily create, simulate and consume immersive experiences as vividly as they can imagine them.
Our Values
- Be Fearless : Unmute yourself. Test boundaries and get proven right.
- Remain Adaptable : Stay comfortable in a continuously changing world. If you’re wrong, concede and move on.
- Educate Your Ego : Selflessly collaborate towards our shared purpose.
About the role
Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining uptime, performance, and operational excellence across compute, storage, and networking environments. Exceptional Linux expertise and advanced automation capabilities are mandatory for success in this role.
What you'll do
- Design, implement, and operate highly available, fault-tolerant infrastructure and services.
- Install, maintain, and upgrade server, storage, and networking hardware in office and colocation facilities.
- Continuously monitor developer and production environments and proactively remediate reliability risks.
- Participate in an on-call rotation and lead incident response efforts, including rapid triage, mitigation, and post-incident root cause analysis.
- Respond effectively under pressure to outages and degradation events to restore service availability.
- Develop, maintain, and continuously improve automation and operational tooling using Bash and Python.
- Partner closely with engineering teams to support development, testing, and production workloads at scale.
Qualifications (required)
- Expert-level Linux systems administration across complex, production environments (this is a core requirement).
- Exceptional proficiency in Bash and Python; advanced scripting and automation skills are mandatory, not optional.
- Proven ability to write maintainable automation and diagnostic tooling for large-scale systems.
- Deep understanding of server hardware, storage subsystems, and datacenter operations.
- Hands‑on experience with virtualization platforms including Proxmox (current), VMware vSphere, and/or OpenShift.
- Strong experience with containerization technologies (Docker, containerd) and orchestration platforms (Kubernetes).
- Experience operating workloads in AWS and/or Microsoft Azure environments.
- Experience implementing observability, monitoring, and alerting using tools such as Prometheus and Grafana.
Additional Qualifications
- Familiarity with systems programming languages such as C, C++, Rust, Go, and/or Julia.
- Relevant certifications such as CompTIA A+, Azure Engineer, or similar are preferred.
- Active government clearance or the ability to obtain one is preferred.
On-Call & Incident Response Expectations
This role includes participation in an on-call rotation supporting developer and production systems. The SRE is expected to respond to incidents outside of normal business hours as required, lead technical incident response efforts, communicate effectively with stakeholders during outages, and produce clear post-incident documentation and corrective action plans.
Compensation Range
$145,000–$165,000 per year (California). This range represents the anticipated base pay for this role; the final offer may vary based on qualifications, experience, and location.
- Medical, Dental, & Vision - 100% covered premiums
- Equity - Stock Options
- 401(k) match
Equal Opportunity Statement
Bolt is committed to building a diverse and inclusive environment in which we recognize and value each other’s differences as well as fostering a culture that promotes its core values: Professionalism, Integrity, and Respect. As an equal opportunity employer, all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, genetic information, national origin, age, disability, or status as a protected veteran.
Location & Sponsorship
Please note that Bolt Graphics does not currently sponsor candidates for this role. This role is strictly based in Sunnyvale, CA and will require someone to be locally based, preferably in the Bay Area.
#J-18808-Ljbffr$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...SuggestedFull timeWorldwide$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SuggestedFull time$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...SuggestedFull time- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...SuggestedFull time
$145k - $175k
...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data...Full timeRemote work- ...Google is seeking a Senior Engineering Manager for Collaboration SRE to lead a multi-site engineering organization across Sunnyvale and Zurich. You will own the... ..., and drive high-impact projects that improve reliability, performance, and scalability of #J-18808-Ljbffr
- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data...
$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide$276.1k - $311.4k
...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a...Permanent employmentFull timeWork at officeWork from home$128k - $216k
...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay... ...a global scale, come make a difference at Fiserv. Sr. Site Reliability Engineer About Clover Clover is a pioneer in the fintech space...Worldwide$160k - $240k
...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay... ...come make a difference at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site Reliability Engineer do at...- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response...Contract work
- ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of... ...are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in...Work experience placementImmediate start
$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of...Full timeContract workWorldwide$132.6k - $214.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps...Full timeWork at officeVisa sponsorshipWork visa- ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python...
$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...$255.7k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation...Full time$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with...Full time$222k - $300.5k
...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps... ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency...WorldwideShift work$262k - $364k
...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with... ...capacity and performance.Build creative engineering solutions to operations and... ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software...- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$217.57k - $260k
...job description explicitly states otherwise, all roles are on-site five days per week at one of our offices in McLean, VA;... ...which can be found here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale infrastructure...Full timeTemporary workWork at officeRemote workFlexible hoursShift work- ...precision that drives great outcomes. Job Summary Key Responsibilities Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching, and career development. Own the reliability, availability, and operational...Full timeWork at officeVisa sponsorshipWork visa
$200k - $260k
...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent...Work at officeHome office$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Sunnyvale, CA
- site reliability engineer Sunnyvale, CA
- IT site lead Sunnyvale, CA
- site safety Sunnyvale, CA
- website content developer Sunnyvale, CA
- site leader Sunnyvale, CA
- on-site clinical research associate (traveling/remote) Sunnyvale, CA
- junior website developer Sunnyvale, CA
- historic site Sunnyvale, CA
- site recruiter Sunnyvale, CA


