Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

3B Staffing LLC

Location: Irving, TX


Duration: 6+ Month Contract to hire


Interview: Onsite. (mandatory)


Term: Hybrid (2 days in office)


Responsibilities

  • Lead architecture and development teams to ensure applications are highly available, reliable, and performant at a global scale.
  • Partner with the architecture team to ensure operability, measurability, and manageability are integrated into business features and enablers.
  • Collaborate with product owners and managers to establish service level objectives (SLOs) for applications and define consequences if objectives are not met.
  • Work with development team members to identify monitoring gaps, improve application performance, and assist with troubleshooting issues.
  • Drive Root Cause Analysis (RCA) of production issues and other failures within the product software, pipeline, or other DevOps support processes or technology.
  • Design, build, and advocate for automated solutions to optimize application/service/platform uptime with minimal human intervention.
  • Participate in an on-call rotation to support troubleshooting and communication efforts outside of normal business hours.
  • Create and implement standards and best practices, driving adoption across development teams and external vendors as applicable.
  • Ensure compliance with all company policies and procedures.
Qualifications
  • Bachelor of Computer Science or related Engineering field required.
  • Master's Degree preferred.
Required Skills
  • 5-7 years of hands-on SRE experience.
  • 1-2 years of leading and mentoring others.
  • Hands-on experience supporting Linux production environments, hands-on administration on Spark, and hands-on experience with MS Azure Cloud technologies.
  • 3-5 years hands-on experience with scripting with bash, perl, ruby, or python required.
  • 3-5 years experience with Docker Datacenter required.
  • 2-4 years of hands-on administration experience on Machine learning platforms required.
  • Minimum of 1 year of experience in Mesos, Kubernetes, OpenShift and/or Deis or other such container/platform-as-a-service orchestrator required.
  • Minimum of 1 year of hands-on experience on CICD tools & Technologies required.
  • Minimum of 1 year of lead experience of site reliability engineering team required.
  • Proven leadership skills and the ability to guide and mentor a team.
  • Strong collaboration and communication skills.
  • A proactive approach to problem-solving and continuous improvement.
  • Passion for automation and operational excellence.
  • Deep expertise in cloud technologies and software development, with a strong technical background.
  • Experience with Java
  • Proficiency in SQL and Powershell.
  • Expertise in defining, implementing, and evaluating Service Level Objectives (SLOs) and Service Level Indicators (SLIs), and associated consequences.
  • Strong skills in performing Root Cause Analysis (RCA) and Problem Management.
  • Extensive experience in cloud native applications Azure/AWS (monitoring, networking, containerization, infrastructure).
  • Proficiency in containerization technologies such as Azure Kubernetes Service, Kubernetes (open source), and Docker.
  • Knowledge of metrics and monitoring tools like Azure Application Insights and Azure Monitor.
  • Familiarity with networking technologies relevant to Azure and AWS, including Azure DNS, Virtual Networks, Azure API Manager, Azure Application Gateway, Akamai WAF/CDN, AWS Route 53, AWS VPC, AWS API Gateway, and AWS CloudFront.
  • Strong experience with Terraform for infrastructure as code.
  • Ability to establish and maintain a culture of learning through the development and sharing of skills, knowledge, processes, and tools; combat traditional silos that create "us and them" environments.
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!