Site Reliability Engineer
3B Staffing LLC
Job Description:
Qualifications:
This position is 60 % SRE and 40% SDE. Also open for candidates to join MTH along with ATL GO team. Must agree to work onsite at either of these two locations based on the team's hybrid schedule. Required Skillset
• Manage and optimize data streaming and API components in OpenShift Onpremise and AWS.
• Proactively review the application's APIs and processes to identify opportunities to optimize the response times for various application components.
• Automate various types of testing including data quality checks, automate delivery to system integration, production and automate deployment for production
• Develop integrations between the application in Onpremise and AWS and our third-party tools (ServiceNow, VersionOne, Sumo)
• Work with teams to create SLI/SLO's .
• Actively monitor and lead troubleshooting of degraded performance and hard to define issues for the platform applications, develop the solution and document artifacts in the back log from root cause analysis.
• Evolve the cloud infrastructure ecosystem for our application suite by experimenting with emerging technologies and completing prototypes to understand benefit
• Design and develop CI/CD pipeline in AWS and Openshift to deploy various application artifacts, including APIs and Data Process Jobs.
• Analyze, design and develop the artifacts to configure the monitoring and alerting metrics so the support engineers can proactively and timely validate, troubleshoot and resolve the issues.
• Maintain data integrity and access control by using AWS security tools and services such as HSM, IAM, etc.
• Understand and develop tools to monitor AWS billing for the services, generate cost related reports and help develop and implement cost optimization strategies.
• Work with enterprise security architects to design and implement data security tools, measures, data encryption, key management; design and develop solutions to address the security vulnerabilities discovered by internal security audit team, as well as by the vendors, security community, etc.; design and develop solutions for support team to regularly scan and review to fix security issues
• Regularly and proactively monitor and analyze the capacity and performance of the platform, work with architecture team to design and implement elastic infrastructure to accommodate the irregular burst of user traffic/requests.
• Work with architecture team to develop backup strategy and implement the backup solution for critical data and application components for service restoration and disaster recovery purpose.
• Work with architecture, infrastructure, and application teams to provide input on continuous improvement on the design, performance and security enhancements. Desired Skillset:
• Deep understanding of the operations of AWS cloud platforms.
• Must be well versed in the automation, scripting, monitoring, including use of tools from the major cloud platforms, including but not limited to OpenShift Cloud Formation, Terraform, Ansible, Shell, Python
• Preferable for candidates with significant technical knowledge with infrastructure layers, including but not limited to: Linux OS, major virtualization platforms, Traditional and software defined network, Load Balancers, firewall, API tools, element/performance/intelligent monitoring tools, storage, backup strategy, etc.
• Significant knowledge and experience in end-to-end operations for enterprise systems and applications, including driving issue resolution for mission critical systems.
• Must have experience working to automate, operationalize and improve the Development/QA using CI/CD tools (Gitlab, Github, Jenkins, Maven, Gradle, Nexus)
• Working experience with Software Release Management. Desired Qualification
• BS degree in Computer Science or a related technical field or equivalent practical experience. Minimum Experience
• 3+ years of related DevOps, SysOps engineering experience with focus on major cloud platforms (AWS preferred).
• 2+ years of application development experience including data streaming, deploying/monitoring high availability critical application components.
• 1+ Years in Site Reliability Engineering organization preferred
• Overall 4-6years of experience
Responsibilities:
As a engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to a broad range of industry leading technologies that will help to drive acceleration. The ideal candidate will have expert design and development capabilities and be positioned to contribute to a growing set of services and features for the ecosystem. This role will be supporting highly available, business critical applications. This role will serve as the escalation point for complex and hard to define issues in both on premise and AWS environments. We are seeking talented engineers, well versed in DevOps technologies, automation, infrastructure orchestration, configuration management, continuous integration, troubleshooting of complex issues, who are not constrained by how "things are usually done".
Qualifications:
This position is 60 % SRE and 40% SDE. Also open for candidates to join MTH along with ATL GO team. Must agree to work onsite at either of these two locations based on the team's hybrid schedule. Required Skillset
• Manage and optimize data streaming and API components in OpenShift Onpremise and AWS.
• Proactively review the application's APIs and processes to identify opportunities to optimize the response times for various application components.
• Automate various types of testing including data quality checks, automate delivery to system integration, production and automate deployment for production
• Develop integrations between the application in Onpremise and AWS and our third-party tools (ServiceNow, VersionOne, Sumo)
• Work with teams to create SLI/SLO's .
• Actively monitor and lead troubleshooting of degraded performance and hard to define issues for the platform applications, develop the solution and document artifacts in the back log from root cause analysis.
• Evolve the cloud infrastructure ecosystem for our application suite by experimenting with emerging technologies and completing prototypes to understand benefit
• Design and develop CI/CD pipeline in AWS and Openshift to deploy various application artifacts, including APIs and Data Process Jobs.
• Analyze, design and develop the artifacts to configure the monitoring and alerting metrics so the support engineers can proactively and timely validate, troubleshoot and resolve the issues.
• Maintain data integrity and access control by using AWS security tools and services such as HSM, IAM, etc.
• Understand and develop tools to monitor AWS billing for the services, generate cost related reports and help develop and implement cost optimization strategies.
• Work with enterprise security architects to design and implement data security tools, measures, data encryption, key management; design and develop solutions to address the security vulnerabilities discovered by internal security audit team, as well as by the vendors, security community, etc.; design and develop solutions for support team to regularly scan and review to fix security issues
• Regularly and proactively monitor and analyze the capacity and performance of the platform, work with architecture team to design and implement elastic infrastructure to accommodate the irregular burst of user traffic/requests.
• Work with architecture team to develop backup strategy and implement the backup solution for critical data and application components for service restoration and disaster recovery purpose.
• Work with architecture, infrastructure, and application teams to provide input on continuous improvement on the design, performance and security enhancements. Desired Skillset:
• Deep understanding of the operations of AWS cloud platforms.
• Must be well versed in the automation, scripting, monitoring, including use of tools from the major cloud platforms, including but not limited to OpenShift Cloud Formation, Terraform, Ansible, Shell, Python
• Preferable for candidates with significant technical knowledge with infrastructure layers, including but not limited to: Linux OS, major virtualization platforms, Traditional and software defined network, Load Balancers, firewall, API tools, element/performance/intelligent monitoring tools, storage, backup strategy, etc.
• Significant knowledge and experience in end-to-end operations for enterprise systems and applications, including driving issue resolution for mission critical systems.
• Must have experience working to automate, operationalize and improve the Development/QA using CI/CD tools (Gitlab, Github, Jenkins, Maven, Gradle, Nexus)
• Working experience with Software Release Management. Desired Qualification
• BS degree in Computer Science or a related technical field or equivalent practical experience. Minimum Experience
• 3+ years of related DevOps, SysOps engineering experience with focus on major cloud platforms (AWS preferred).
• 2+ years of application development experience including data streaming, deploying/monitoring high availability critical application components.
• 1+ Years in Site Reliability Engineering organization preferred
• Overall 4-6years of experience
Responsibilities:
As a engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to a broad range of industry leading technologies that will help to drive acceleration. The ideal candidate will have expert design and development capabilities and be positioned to contribute to a growing set of services and features for the ecosystem. This role will be supporting highly available, business critical applications. This role will serve as the escalation point for complex and hard to define issues in both on premise and AWS environments. We are seeking talented engineers, well versed in DevOps technologies, automation, infrastructure orchestration, configuration management, continuous integration, troubleshooting of complex issues, who are not constrained by how "things are usually done".
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Atlanta, GA vacancy
- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer...SuggestedFull timeWorldwideFlexible hours
- Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence...SuggestedWorldwide
$35 - $45 per hour
DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'...SuggestedRemote work$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours$178.13k - $205.4k
...Qualification Bachelor's degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)... ...websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a...SuggestedWork at officeRemote workFlexible hours- ...Technical Support Specialist In Site Reliability Engineering (Sre) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management. The...
$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and scalable...$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office- ...Job title: Site Reliability Engineer (SRE) Location: Atlanta, GA 30303 Duration of the project: 12 Months Strong expertise in Ansible with an SRE background Ability to review and test GitLab Duo generated code CICD pipelines (GitLab, GitHub Actions) Infrastructure...
- ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to...
- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$81.1k - $187k
...operational improvement. You will work alongside experienced engineers, partner teams, and vendors to maintain and optimize the infrastructure... ...solutions that reduce recurring operational work and improve reliability at scale. Responsibilities NOC Operations Use...Temporary workWorldwideFlexible hoursShift workNight shift- ...availability. • Automation Experience with Build/deployment, Software Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in automating manual processes using Python, Ruby, Unix Shell (bash, ksh...Immediate start
$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested...Contract workLocal areaImmediate start$70 - $85 per hour
...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,...Temporary workLocal areaFlexible hours3 days per week- ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it...Full timeLive inWork at office
- ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's... ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps...Work at officeLocal areaRemote workWork from homeWorldwide
- Kforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and optimizing...Temporary workRemote work
- ...build, and test of the company's first combined turbojet-ramjet engine and is now being scaled through its first flight vehicle... ...capabilities to the warfighter. Hermeus is seeking a Senior Software or Site Reliability Engineer to join the Information Team and take charge of...Full timeRemote work
- Direct message the job poster from STAFFWORXS Delivery Manager @ STAFFWORXS | US IT Recruitment Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site Reliability Engineer (SRE) to join our team in Atlanta, GA. This hybrid role offers the opportunity to work...Contract work
$71.6k - $119.4k
...deployment support, and security improvements. You'll help implement automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You'll gain exposure to a wide range of cloud technologies, automation tools, and data...Temporary workInternshipLocal area- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
- ...Consultancy and Information Technology Enabled Services.Job DescriptionSCM System EngineerSCM Continuous Integration / Delivery Build Team Engineer with experience in Application Service and Web Application Build, Deployment and Release Management and experience in establishing...Permanent employmentFull timeH1b
- OneTrust is seeking a Senior Software Engineer in Atlanta, Georgia. The role involves designing and maintaining a reliable application platform, collaborating with engineering teams, and enhancing customer experiences through observability tools. The ideal candidate will...
$101.5k - $169.1k
...include an incentive program.Job DescriptionThe Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train... ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures. Maintain...Full timeWork at officeRemote workVisa sponsorshipFlexible hours$141.3k - $237.4k
...AT&T, you won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of...Full timeTemporary workWork at officeLocal areaRelocation- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$143k - $191k
...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems...Full timeWork experience placementImmediate start- ...YesApply: JobGeorgia-Pacific's Corrugated Division is seeking a Reliability Engineer to support our growing Mailers network, including facilities... ...programs, and driving standardization across multiple sites. Success in this position requires strong technical capability...For contractorsRemote workWorldwideVisa sponsorshipFlexible hours
- ...coaches. That’s how we’re UNSTOPPABLE for our employees!Are you ready for the next chapter in your Uncarrier journey? The System Reliability Engineer (SRE) improves and protects the software and systems behind all of T-Mobile's IT services, including management of...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- junior website developer Atlanta, GA
- website content developer Atlanta, GA
- on site coordinator Atlanta, GA
- website coordinator Atlanta, GA
- site leader Atlanta, GA
- site recruiter Atlanta, GA
- historic site Atlanta, GA
- on-site clinical research associate (traveling/remote) Atlanta, GA


