Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$100k - $180k

Bright Vision Technologies

Site Reliability Engineer (SRE) - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Site Reliability Engineer (SRE)
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000-$180,000 Annually
Experience Required: 10+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary:
We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.
Key Responsibilities
  • Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services, and use those measures to drive concrete engineering and prioritization decisions.
  • Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed, and ensuring high-quality post-incident reviews that drive lasting improvements.
  • Design and implement comprehensive monitoring, logging, and tracing strategies using Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar tooling so that operators have rich, actionable visibility into system behavior.
  • Build and maintain robust on-call processes, runbooks, and escalation paths that reduce mean time to detect and mean time to resolve while protecting the well-being of the engineers on rotation.
  • Automate operational toil aggressively by writing production-grade tooling in Python, Go, Bash, or similar languages, replacing manual workflows with reliable, auditable automation.
  • Architect and operate large-scale Kubernetes clusters and container-based workloads, including autoscaling, capacity planning, network policy, and integration with service meshes.
  • Design CI/CD pipelines that promote safe, frequent, and observable releases, supported by automated testing, canary deployments, feature flags, and progressive rollout strategies.
  • Lead capacity planning and performance engineering activities, building models that predict growth and stress, and validating those models through load testing and chaos experiments.
  • Partner closely with application development teams to embed reliability practices early in design - including failure-mode analyses, graceful degradation patterns, and dependency hardening.
  • Strengthen the platform's resiliency through chaos engineering, fault injection, dependency isolation, retries, timeouts, circuit breakers, and well-tested failover paths.
  • Drive continuous improvement of security posture in collaboration with security teams, including patch management, vulnerability remediation, and secure-by-default platform defaults.
  • Contribute to the technical roadmap for reliability tooling, observability platforms, and developer-experience improvements that reduce friction and improve outcomes for engineering teams.
  • Mentor engineers across the organization on SRE practices and foster a strong, blameless culture of operational excellence.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
  • Ten or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.
Preferred Qualifications
  • Experience defining and operationalizing SLOs and error budgets in real production environments.
  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Background in capacity planning, performance engineering, or large-scale load testing.
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul.

How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at View phone number on click.appcast.io. Learn more about Bright Vision Technologies at
Bright Vision Technologies is an Equal Opportunity Employer.


Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in United States vacancy
  •  ...Primary Location: Louisville, Kentucky V-Soft Consulting is currently hiring for a SRE-Site Reliability Engineer for our premier client in Louisville, Kentucky . Education And Experience » ~6+ years in IT infrastructure, with 3+ years focused on site reliability... 
    Suggested
    Permanent employment
    Full time
    Work experience placement
    Currently hiring
    Local area

    V-Soft Consulting Group, Inc.

    Louisville, KY
    1 day ago
  • $70.8k - $131.4k

    Job DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate...  ...review is needed.Key ResponsibilitiesSupport and maintain SRE operational tooling, including dashboards, alerts, runbooks... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    Saint Paul, MN
    3 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Suggested
    Contract work

    2T Consulting

    Atlanta, GA
    a month ago
  •  ...prioritize a diverse F5 community where each individual can thrive.Role SummaryWe are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-... 
    Suggested
    Full time
    Local area
    Immediate start
    Shift work
    Night shift
    Afternoon shift
    Weekday work

    F5 Networks

    Reston, VA
    20 hours ago
  • $135k - $155k

     ...buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide the reliability and performance... 
    Suggested
    Flexible hours

    Thomas

    Denver, CO
    2 days ago
  • $175k - $215k

     ...experiences — and we’re constantly looking for new ways to enhance these exciting experiences.Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational priorities and functional... 

    Disney Interactive

    Orlando, FL
    5 days ago
  •  ...the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability...  ...teams, you will apply Site Reliability Engineering (SRE) principles to improve system availability,... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    4 days ago
  •  ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Permanent employment
    Full time
    Work experience placement
    Local area

    Scientific Games Corporation

    Alpharetta, GA
    3 days ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident response...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    5 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    3 days ago
  •  ...Period of performance: Up to 2 years in duration MUST HAVES: Minimum of 8 years of experience as a Site Reliability Engineer with a strong understanding of SRE principles for highly scalable and reliable systems Possess a bachelor's degree Experience working... 
    Local area
    Relocation package
    3 days per week

    Beyond SOF

    Vienna, VA
    5 days ago
  •  ...We’re seeking a highly skilled Site Reliability Engineer (SRE) to join our engineering team and help ensure the reliability, scalability, and performance of our systems. As an SRE, you’ll blend software engineering with systems engineering to build and maintain resilient... 
    Temporary work
    Interim role
    Remote work
    Flexible hours

    OutSolve - Beyond Compliance

    Mission, KS
    2 days ago
  •  ...Description The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group's proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure... 
    Work at office
    Remote work

    CenCore

    United States
    3 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    2 days ago
  • $142.3k - $263.3k

     ...recently the new LIDAR iPad sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the next level. In this role,...  ...Computer Vision Organization. As a main contributor to our SRE team you will develop and maintain infrastructure, tooling,... 
    Work experience placement
    Relocation

    Apple

    San Diego, CA
    3 days ago
  • $165k - $225k

     ...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure... 
    Remote work
    Flexible hours

    Moonlite AI

    Chicago, IL
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: North Little Rock AR (onsite) Duration: Contract Required/Desired Skills: • Strong web development skills with a strong focus in C#/.NET • Someone who currently works in a hybrid skillset of BOTH.Net development AND... 
    Contract work

    Software Technology Inc

    North Little Rock, AR
    1 day ago
  •  ...Site Reliability Engineer (SRE) - Security Infrastructure Position Summary We are seeking an SRE to support reliability, scalability, and operational excellence for a large-scale network security transformation initiative. This role will focus on monitoring... 

    IS3 Solutions

    Cary, NC
    4 days ago
  •  ...Site Reliability Engineer (SRE) Immediate need for a talented Site Reliability Engineer (SRE). This is a 12+ months contract opportunity with long-term potential and is in Chicago, IL (Hybrid). Key Requirements and Technology Experience: ~ Must have skills:... 
    Contract work
    Local area
    Immediate start

    Pyramid Corporation

    Chicago, IL
    1 day ago
  •  ...globe. Join us on this journey to redefine resource management—and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Temporary work
    Worldwide

    airapps

    San Francisco, CA
    23 hours ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system... 

    Mybridge

    Seattle, WA
    2 days ago
  •  ...Site Reliability Engineer (SRE) We are seeking an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of enterprise applications and infrastructure. The ideal candidate will have strong expertise in production support... 

    AceStack LLC

    Strongsville, OH
    1 day ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Iselin, NJ
    1 day ago
  •  ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity... 

    Methodic

    San Francisco, CA
    2 days ago
  •  ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading applications and Fidessa flows used by internal business users and external clients. This role is responsible for ensuring... 
    Permanent employment
    Work at office
    Afternoon shift

    RIT Solutions, Inc.

    New York, NY
    3 days ago
  • $110k - $120k

     ...documentation that you are a U.S. citizen to qualify. Summary: The Site Reliability Engineer owns the day-to-day health, security, and usability of the organization's ICAM environment while applying SRE practices to keep delivery platforms reliable, observable, and... 
    Contract work
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    Cherokee Federal

    Boulder, CO
    4 days ago
  •  ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with... 
    Contract work

    AceStack LLC

    Brentwood, TN
    1 day ago
  •  ...We are seeking a Site Reliability Engineer (SRE) to improve the reliability, scalability, and performance of cloud-based applications. The ideal candidate will automate infrastructure, monitor production systems, troubleshoot incidents, and collaborate with software engineering... 

    Mybridge

    Austin, TX
    2 days ago
  •  ...Sr. Site Reliability Engineer (SRE) New York City, NY - LOCALS ONLY Hybrid, 3 days 6-Month Contract Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment... 
    Contract work
    Local area

    RIT Solutions

    New York, NY
    5 days ago
  • $40 - $80 per hour

     ...Job Title: Site Reliability Engineer (SRE) Duration (Contract): 12 Months Client Location: Southlake, TX Location Preference: Onsite Job Description: As a Site Reliability Engineer (SRE) , you will be responsible for improving the reliability, scalability... 
    Hourly pay
    Contract work

    Smart IMS

    Southlake, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!