Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- Site Reliability Engineering (SRE)

$106.5k - $177.5k

Noctua Technology

The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don’t just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations.

We are seeking a motivated Site Reliability Engineer (SRE) to join our dynamic team. As a key contributor, you will apply software engineering principles to operations, focusing on the reliability, scalability, and performance of production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure.

  • Location : Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams.
  • Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.

Key Responsibilities

Site Reliability Engineering

  • Define, measure, and report on Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure system reliability and uptime.
  • Develop and deploy Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools, with an emphasis on repeatability and change management.
  • Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.
  • Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.

Toil Reduction and Incident Management 

  • Implement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.
  • Eliminate toil by extensively automating operational tasks, including provisioning, patching, and deployments, using scripting and configuration management tools such as Python, Bash, or Ansible.
  • Conduct post-incident reviews (blameless postmortems) to drive continuous improvement in system reliability and operational processes.

Testing and Service Resiliency

  • Implement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.
  • Proactively identify and address system weaknesses and ensure performance under stress.
  • Support disaster recovery and high availability strategies through backup and failover planning.

Collaboration and Knowledge Sharing

  • Collaborate with development teams to improve the operability and production readiness of applications from design through deployment.
  • Create and maintain documentation for cloud architectures, deployment processes, and best practices.
  • Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.

Stakeholder Communication

  • Provide technical guidance and support to clients and internal teams on cloud infrastructure and reliability best practices, with a focus on defining Service Level Agreements (SLAs).
  • Act on client feedback to refine and enhance cloud solutions.
  • Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.

Continuous Learning and Innovation

  • Stay updated on the latest developments in cloud infrastructure and technology trends.
  • Drive innovation by proposing and implementing new techniques and technologies.

Qualifications

  • 1-5 years of experience in site reliability engineering, cloud engineering, or related fields.
  • Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.
  • Proficiency in Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
  • Experience with containerization and orchestration tools like Docker and Kubernetes.
  • Knowledge of networking concepts, cloud security best practices, and identity management.
  • Experience with programming or scripting languages such as Python, Bash, or Go.
  • Familiarity with CI/CD pipelines and DevOps methodologies.
  • Strong problem-solving skills and the ability to troubleshoot complex cloud environments.
  • Effective communication skills and a willingness to learn and collaborate.

Preferred qualifications:

  • Bachelor's or advanced degree in Computer Science or a related field.
  • Any of the below cloud certifications:
    • Google Cloud Professional Cloud Architect
    • Google Cloud Professional Cloud DevOps Engineer
    • AWS Certified Solutions Architect
    • AWS Certified Developer
    • AWS Certified SysOps Administrator
    • Azure Solutions Architect Expert
  • CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.

Salary Range : $106,500 - $177,500

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer- Site Reliability Engineering (SRE) in United States vacancy
  • $70.8k - $131.4k

     ...DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and...  ...ResponsibilitiesSupport and maintain SRE operational tooling, including...  ...systems engineering, production operations, software engineering, or a related technical... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    Eagan, MN
    4 days ago
  • $101k - $161k

     ...artificial intelligence, and software-defined networking to...  ...prestigious awards, such as Best Engineering Team, Best Company for Diversity...  ...Work WithWe’re looking for Site Reliability Engineers to join our...  ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine... 
    Suggested

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $135k - $155k

     ...global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual...  ...and performance of our infrastructure and software systems across several engineering teams and influence... 
    Suggested
    Flexible hours

    Thomas

    Denver, CO
    3 days ago
  •  ...Opportunity: We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We...  ...-driven approaches to observability and reliability. What you’ll do: • Evangelize SRE mindset and solve problems through systematization... 
    Suggested

    Mindlance

    Austin, TX
    4 days ago
  •  ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Suggested
    Permanent employment
    Full time
    Work experience placement
    Local area

    Scientific Games Corporation

    Alpharetta, GA
    17 hours ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    4 days ago
  •  ...Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure...  ..., and who treats infrastructure like software. You'll have significant ownership...  ...What You Bring: ~​​6+ years in SRE, DevOps, or infrastructure engineering... 
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    17 hours ago
  • $207k - $300k

     ...design consulting, developing software platforms and frameworks,...  ...for changes that improve reliability and velocity.Practice...  ...in Computer Science or Engineering.Experience mentoring engineers...  ...cross-functional teams.Site Reliability Engineering (SRE) combines software and... 

    Google

    New York, NY
    4 days ago
  • $172k - $300k

     ...Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable,...  ..., and operate autonomous-vehicle software.As one of our founding SREs, you will...  ...depending on heroics.If you are an expert in SRE practices who loves building the... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    2 days ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    17 hours ago
  • $160k - $185k

     ...And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement...  ...support. This role operates at the intersection of software engineering and infrastructure, driving automation, reducing... 
    Work at office
    Local area
    Remote work
    Work from home

    Planet Fitness

    Hampton, NH
    1 day ago
  •  ...Job Title:  Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT  Location:  Remote Citizenship Requirement...  ...Role/Responsibilities The Site Reliability Engineer (SRE) for Azure Government & Infrastructure plays a critical role... 
    Full time
    Remote work
    Monday to Friday

    Quzara LLC

    United States
    1 day ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected... 
    Full time
    H1b
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    5 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    1 day ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 

    2T Consulting

    Decatur, GA
    3 days ago
  •  ...automate, deploy, and operate highly reliable cloud systems supporting mission-critical...  ...role is centered on DevSecOps and site reliability engineering, with a strong emphasis on deployment...  ...years of professional experience as an SRE, DevOps, reliability, infrastructure,... 
    Permanent employment
    Remote work

    Quindar

    United States
    3 days ago
  • $142.3k - $263.3k

     ...San Diego, California, United States Software and Services The Video Computer Vision...  ...sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the...  ...Organization. As a main contributor to our SRE team you will develop and maintain... 
    Work experience placement
    Relocation

    Apple

    San Diego, CA
    17 hours ago
  •  ...management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the...  ...of our systems. You will work at the intersection of software development and operations, implementing automation,... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    5 days ago
  • $73.41 per hour

     ...Number Of Positions 1 Job Description Job Title: Site Reliability Engineer (SRE) Location: Pennington, NJ / Charlotte, NC Duration:...  ...present positions to business stakeholders. Background in software development or DevOps. Preferred Skills: Experience... 
    Hourly pay
    Contract work

    Bucher & Christian Consulting Inc.

    Pennington, NJ
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    Los Angeles, CA
    2 days ago
  •  ...leading provider of real estate closing and title insurance software. A division of Fidelity National Financial (NYSE: FNF),...  ...we looking for? SoftPro is seeking a well-rounded Site Reliability Engineer (SRE) to join our Cloud Operations Team in our Raleigh, NC... 
    Hourly pay
    Work at office
    Remote work

    Fidelity National Financial

    Raleigh, NC
    3 days ago
  •  ...Site Reliability Engineer (SRE) We are seeking an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of enterprise applications and infrastructure. The ideal candidate will have strong expertise in production support... 

    AceStack LLC

    Phoenix, AZ
    3 days ago
  •  ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering...  ...application health ~ Understanding of SRE principles, including observability,... 
    Remote work

    Anveta

    United States
    3 days ago
  •  ...Job Title A background in SRE or DevOps in AVD Experience supporting large organizations in globally diverse locations Knowledge and experience of supporting and administrating enterprise level Windows devices and configuration and management systems At least three... 

    Omni Inclusive

    Jacksonville, FL
    3 days ago
  •  ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with... 
    Contract work

    AceStack LLC

    Brentwood, TN
    3 days ago
  • $111.61k - $131.3k

     ...support, product/project management, or application developmentPreferred Skills/ExperienceStrongexpertiseinSiteReliabilityEngineering(SRE),DevOps,ProductionSupport,PlatformEngineering,andDistributedSystemsOperations.Experienceleadingtechnicalteams,incidentresponseefforts... 
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Atlanta, GA
    2 days ago
  • $119.8k - $234.7k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type:...  ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s...  ...organization. As a Senior Linux SRE (Site Reliability Engineer), your main focus will be to... 
    Ongoing contract
    Permanent employment
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Hillsboro, OR
    2 days ago
  • Vice President - Site Reliability Engineering (SRE) - The Core EngineeringWHAT WE DOSite Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly... 

    Goldman Sachs

    New York, NY
    4 days ago
  •  ...markets and shape the future of our communities.This is a Software Engineering position at Director level, which is part of the job...  .... This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and... 

    Morgan Stanley

    Alpharetta, GA
    1 day ago
  • $63 - $90 per hour

     ...Job Summary Our client, a leader financial services provider, is seeking a Senior DevOps Engineer / Site Reliability Engineer (SRE) to support enterprise backup, cyber recovery, and platform resiliency initiatives. The ideal candidate brings experience in backup and... 
    Permanent employment
    Contract work
    Local area

    KellyMitchell Group

    Chicago, IL
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- Site Reliability Engineering (SRE). Be the first to apply!