Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

Jobless

About the team

Sight Machine is built on the shoulders of a unique, robust and highly scalable Infrastructure as Code model. This enables the creation and operation of customer instances in our ecosystem in a standardized and simplified manner. We are looking for team members to help us build, maintain, and improve the infrastructure that makes Sight Machine the leading provider of Manufacturing Data Pipelines and Analytics.

Great things happen when people can bring their authentic selves to work. We empower all of our team members to share their perspectives, passions and experiences because collectively we make a better, stronger team through always “open communications” mind.

Our team collaborates closely with peers & cross functional stakeholders throughout the business, our clients on the forefront of digital transformation, and the cutting edge of digital manufacturing thought leadership.

Sight Machine has offices in San Francisco, CA and Ann Arbor, Mi. We do have a remote-friendly culture with people based all around the US and the rest of the world. For this role in particular, the ideal candidate is located near either of our offices and willing to work in a hybrid capacity. We would still consider 100% remote for exceptional candidates if they aren’t located near an office.

About the role

Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight Machine's platform. You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM gateways, agent orchestration, and the operational patterns that come with non-deterministic workloads.

This is a senior level IC role. You'll help set and drive technical direction for infrastructure and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems problems while still being hands-on with code, infrastructure, and incidents.

Success requires deep technical range, sound judgment on risk vs. customer impact, and the ability to influence architecture decisions across Development Engineering without formal authority.

What You’ll Actually Work On
  • Champion an agentic-AI-first engineering mindset: identify where AI-driven automation and agent-based tooling can replace manual toil, and hold that work to the same quality, testing, and reliability bar as any other production system
  • Evolve reliability practices for meeting reliability SLO’s, error budgets, drive incident postmortems to systemic (not just symptomatic) fixes, and lead reliability reviews for new services before they hit production
  • Troubleshoot and resolve the org's most complex, cross-layer systems problems CI/CD, container orchestration, networking, OS, cloud resources, databases, and increasingly, agentic AI/LLM orchestration layers
  • Design, build, and operate the infrastructure supporting agentic AI workloads, LLM gateway routing, agent orchestration frameworks, monitoring of non-deterministic/AI-driven services, and the operational tooling needed to run them reliably at scale
  • Architect and instrument monitoring, alerting, and observability infrastructure for critical services, with an eye toward what "critical" means for AI-driven systems specifically
  • Author and continuously improve operational runbooks and automation, increasingly incorporating agentic/AI-assisted tooling (e.g., automated triage, AI-assisted incident response) where it measurably reduces toil
  • Design and build internal platforms and developer tooling that other engineers build on top of
  • Participate in on-call coverage and help evolve the program as we scale including escalation paths and reducing avoidable pages through better automation
  • Bring a startup mindset of daily engagement: staying close to what's breaking, what customers are hitting, and where the team needs help, even outside a formal ticket or rotation
  • Mentor senior and mid-level engineers; act as a technical sounding board across teams
  • Proactively identify and drive cross-team initiatives that improve stability, reliability, and availability, this is expected to be self-directed, not assigned
What We’re Looking For
  • Demonstrated experience designing, building, or operating agentic AI/LLM-based systems in production, held to the same quality-first, test-driven rigor as traditional infrastructure code, not just prototype-grade work
  • Embody a quality-first and security-first culture in all that you do
  • 10+ years of experience with Kubernetes/Docker in at least one top-tier cloud provider (Azure, GCP, AWS), including production-scale multi-tenant or multi-cluster environments
  • 10+ years coding experience (Python, Go, Java, or similar) with a track record of building tools/platforms used by other engineers, not just scripts
  • 10+ years with IaC and CI/CD tooling (Terraform/OpenTofu, FluxCD or similar GitOps tooling, Jenkins/GitHub Actions)
  • Strong, provable Linux and networking fundamentals (TCP/IP and application-layer)
  • Practical experience integrating or operating LLM/agentic AI systems in a production context this can be API-based orchestration, LLM gateways, or agent frameworks
  • A track record of authoring technical documentation (design docs, ADRs, runbooks) that other engineers actually use
  • Demonstrated mentorship of other engineers, without needing formal management authority to do it
  • Strong bias for action over endless planning, hands-on, has made mistakes, learned from them, and can weigh risk vs. customer impact under pressure
  • Clear, empathetic communicator, comfortable pushing back on architecture decisions across teams
  • Operational experience with monitoring/alerting systems (Prometheus, Grafana, Loki, Sentry, Signoz or equivalents)
  • Deep understanding of cloud performance, able to diagnose and resolve bottlenecks others can't
Nice to Have
  • Experience with elements of our current tech stack are a plus: Kubernetes, FluxCD, Terraform, Helm Charts, Prometheus, Elasticsearch, Python, Java, Kafka, Postgres, and Jenkins
  • Previous experience or a keen interest in industrial IoT, analytics, or manufacturing a plus
Team Culture

Great things happen when people can bring their authentic selves to work. We empower all of our employees to share their perspectives, passions and experiences because collectively we make a better, stronger team. Our team members collaborate closely with peers & cross functional stakeholders throughout the business, our clients on the forefront of digital transformation, and the cutting edge of digital manufacturing thought leadership.

We take pride in our self-starter culture where employees are enabled and encouraged to achieve their professional goals through leadership guidance, learning and development. Our philosophy is that careers are continuous journeys, and we dedicate time and offer resources so that employees can reach their full potential.

Benefits + Perks

We value you at and outside of work and know your loved ones are important. Our benefits are designed to support you and your family’s health through life’s expected and unexpected events.

Our Benefits Include:

  • Competitive Salary + Stock Options
  • Health Care Coverage + Life Insurance + Health Savings Account + Flexible Spending
  • Account (includes spouse + children)
  • Flexible Vacation Policy
  • Adaptable Working Schedule and Environment
  • Our Perks Include:
  • Casual Dress Attire
  • Hybrid work flexibility
  • Catered Lunches, Snacks and Beverages
  • Commuter Savings Program
  • Company Outings
  • Designated Volunteering Hours + Group Volunteer Events

Sight Machine is proud to be an equal opportunity employer and considers candidates regardless of age, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. Sight Machine also considers qualified applicants regardless of criminal histories, consistent with legal requirements.

About Sight Machine, Inc.

Sight Machine strengthens manufacturers by providing the industry’s only standard data model and system-level visualization capabilities. By integrating all crucial data into a single innovative platform, everyone involved in the fabrication process can visualize, contextualize and examine data in one intuitive interface.

Sight Machine is committed and mission-driven to improve lives, strengthen communities and make the world cleaner through continuously re-envisioning manufacturing processes - making them more efficient, sustainable and absolute.

Founded in Michigan in 2011 and expanded to San Francisco in 2012, Sight Machine blends the spirit of technology innovation and the down to earth style of Detroit manufacturing. Our team includes early leadership from Yahoo, Tesla Motors and Oracle. Together, we share wide industry knowledge and a commitment to advance manufacturing to a more sustainable future.

We take pride in our self-starter culture where employees are enabled and encouraged to achieve their professional goals through leadership guidance, learning and development. Our philosophy is that careers are continuous journeys, and we dedicate time and offer resources so that employees can reach their full potential.

Sight Machine is proud to be an equal opportunity employer and considers candidates legally authorized to work in the US regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. Sight Machine also considers qualified applicants regardless of criminal histories, consistent with legal requirements.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Eastern, KY vacancy
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    5 days ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Suggested
    Work experience placement

    Socket

    Eastern, KY
    1 day ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Suggested
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    3 days ago
  • $175k - $220k

     ...reach is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role... 
    Suggested

    Vertafore Career Center

    Eastern, KY
    3 days ago
  • $169.3k - $304.7k

     ...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Akamai

    Eastern, KY
    3 days ago
  • $194k - $237k

    ## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026...  ...sponsorship.**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability and... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Eastern, KY
    1 day ago
  • $178k - $300k

     ...will make an Impact (Job Summary) SPX is adiverse team of unique individuals who all make an impact. As a Senior Software System Engineer you will actively participate in the development of radio frequency signal acquisition, geolocation and analysis systems used for... 
    Permanent employment
    Work experience placement
    Work at office
    Worldwide
    Long distance
    Flexible hours

    TCI International Inc

    Eastern, KY
    3 days ago
  • $350k

     ...We’re looking for generalist infrastructure and systems engineers to help build the systems that power our foundation models and...  ...models and build the underlying infrastructure for the clusters to reliably and safely train frontier models. Examples might include building... 
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Thinking Machines Lab Inc.

    Eastern, KY
    1 day ago
  • $216k - $270k

     ...competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and...  .... About Us: At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products... 
    Full time
    Live in

    AI Chopping Block

    Eastern, KY
    1 day ago
  •  ...Senior Software Engineer, Platform FullTime Professional Round Rock, TX, US 16 days ago Requisition ID: 1024 Job Purpose/Summary...  ...involves turning the approved architecture and data model into reliable production services spanning object storage, hot DBs, and data... 
    Full time
    Contract work

    Evolving Solution Services

    Eastern, KY
    5 days ago
  • $145k - $198k

     ...AI C3 AI is looking for a highly motivated Senior Software Engineer - Platform to join our Platform Engineering team. You will play...  ...to provision, manage, and improve cloud environments reliably and consistently. Build and maintain internal tools and services... 

    C3 AI

    Eastern, KY
    2 days ago
  • ## Platform Engineering InternApply: Cedar Rapids, IA: Part time: Posted 2 Days Ago: JR1240GreatAmerica is a highly successful entrepreneurial...  ...technologies, infrastructure, and automation, and in the reliable operation of enterprise technology platforms. GreatAmerica's... 
    Full time
    Part time
    Bank staff
    Internship
    Work at office

    GreatAmerica Financial Services

    Eastern, KY
    5 days ago
  • $208k - $312k

     ...Engineering New York City, San Francisco Full Time About Vercel: Vercel is the agentic infrastructure company. We free people...  ..., ship, and operate our systems. Core Platform owns Vercel’s reliability, resilience, and engineering velocity. We’re looking for... 
    Full time
    Work at office
    Remote work
    Work from home
    Worldwide
    Monday to Friday
    Flexible hours

    Cacheflow

    Eastern, KY
    1 day ago
  • $204k - $247k

     ...Senior Software Engineer Join us in building the next generation of AI infrastructure that will power innovation across our organization. We're seeking a senior software engineer to support our Platform team, which owns the AWS infrastructure, Kubernetes clusters,... 

    Neuralsolutions

    Eastern, KY
    5 days ago
  • $180k - $250k

     ...by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on. You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling... 
    Temporary work
    Local area
    Relocation package

    Falai

    Eastern, KY
    1 day ago
  • $142.3k - $263.3k

     ...and help us leave the world better than we found it.Core Data and Intelligence (CDI) is the central data, AI, and machine learning engineering team in ASE Media. We own the primitives other teams build on: privacy-compliant data collection, canonical datasets, AI and... 
    Contract work
    Work at office
    Relocation

    Apple

    Eastern, KY
    1 day ago
  •  ...while some have prior security experience, many have been successful at Vanta without it. We are looking for a Senior Software Engineer to join our Product Platform org - the group that powers product development across Vanta by building the shared foundations that... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Apply

    Eastern, KY
    5 days ago
  • $92.5k - $209.5k

     ...facing CI/CD capabilities and is focused on building secure, high-performance, enterprise-grade DevOps solutions. We're looking for engineers who are passionate about distributed systems, cloud infrastructure, and solving complex technical challenges. This is an... 
    Temporary work
    Flexible hours

    ORACLE Deutschland B.V. & Co. KG

    Eastern, KY
    5 days ago
  • $220k - $300k

     ...Representative Projects Context Engineering & Agent Infrastructure. Build the platform...  ...least 1+ year focused on AI/ML engineering. Staff candidates will typically have 8+ years...  ...challenges of making autonomous systems reliable when the stakes are real. A bias... 

    Harvey

    Eastern, KY
    2 days ago
  • $200k - $250k

     ...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to...  ...it out on their hardest use case, and get agents running reliably at scale. This is a hands-on, highly technical team. Solutions... 
    Work at office
    Flexible hours

    LangChain, Inc.

    Eastern, KY
    1 day ago
  • $209.1k - $275.1k

     ...Enterprise Premier Segment desires to bring onboard a Solutions Engineer (SE), that will be working as a trusted advisor and partner to top...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to... 
    Full time
    Temporary work
    Local area
    Immediate start
    Flexible hours

    Cisco

    Eastern, KY
    3 days ago
  • $176.1k - $308.2k

     ...It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That...  ...discipline. About the role We are looking for a Staff FinOps AI GovernanceLeadto drive financial accountability and optimization... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    SmartRecruiters

    Eastern, KY
    16 hours ago
  •  ...system interfaces, dependencies, and technical requirements across integrated systems. Coordinate integration activities among engineering, cybersecurity, cloud, data, operations, and mission-partner teams. Identify and mitigate technical risks and integration... 
    Flexible hours

    International Executive Service Corps

    Eastern, KY
    1 day ago
  • $228.5k

     ...mission-driven Senior Manager, Platform Engineering who believes that great products emerge...  ..., data-driven systems that need to work reliably during critical moments, scale as our partners...  ...a culture of care to ensure that our staff are best equipped to lead happy, healthy... 
    Full time
    Temporary work
    Remote work
    Home office
    Flexible hours
    Shift work

    murmuration

    Eastern, KY
    1 day ago
  •  ...is known for pushing the boundaries of innovation, redefining engineering capabilities, and driving advances in various sciences through...  ...multidisciplinary engineering teams.**Key Responsibilities:**•Site Hardware deployment, integration, checkout and test through verification... 
    Hourly pay
    Contract work

    Apex Systems

    Eastern, KY
    3 days ago
  • $117.2k - $313.7k

     ...Distributed Systems Software Engineer - Public Cloud (Mid/Senior/Lead/Principal) Distributed...  ...users count on our platform to be highly reliable, lightning fast, supremely secure, and to...  .... You have experience balancing live-site management, feature delivery, and retirement... 

    Salesforce.Com Inc

    Eastern, KY
    1 day ago
  • $166.9k - $230.9k

     ...with purpose, and motivated by work that truly matters, we’d love to hear from you. The Team: As a member of our Servicing Engineering team, you will play a pivotal role in ensuring smooth loan management, optimizing collections, and cultivating strong customer relationships... 
    Summer work
    Currently hiring
    Local area
    Remote work
    Work from home

    OhioX

    Eastern, KY
    5 days ago
  • $195k - $220k

     ...computational breakthroughs. Join a world-class team of scientists, engineers, and business professionals to advance the state-of-the-art in...  ...and cloud platforms, driving key trade-off decisions across reliability, security, scalability, and cost. This position requires solid... 
    Temporary work
    Work at office
    3 days per week

    Socket

    Eastern, KY
    1 day ago
  • $150k - $200k

     ...The Director, Platform Engineering, is the senior engineering leader responsible for the architecture, deliver, operations, reliability, and scalability of the company's enterprise Data Lakehouse platform. This role leads teams responsible for platform engineering and... 
    Full time

    Dynata

    Eastern, KY
    5 days ago
  •  ...This is a key role in a small and successful R&D team, building tools and engines used by the biggest names in media streaming. It's a critical, visible position where you can have a genuine impact on the end products which are listened to by millions. It's as much... 
    Remote work

    The Audio Programmer Ltd

    Eastern, KY
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!