Site Reliability Engineer
ARANGO INC
Site Reliability Engineer
At ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments.
As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams.
Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!
Key responsibilities include:
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
- Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
- Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
- Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
- Develop strategies for disaster recovery, high availability, and fault tolerance.
- Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
- Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
- Participate in on-call rotations to support critical production systems and respond to incidents.
- Collaborate with cross-functional teams to improve overall system reliability and scalability.
- Collaborate with the Customer Success team to resolve customer issues.
Required skills and qualifications:
- Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
- Proficiency with Kubernetes in managing large-scale, distributed systems.
- Experience with cloud providers such as AWS and Google Cloud (GCP).
- Solid understanding of networking, security practices, and troubleshooting methods.
- Understanding of Linux internals (processes, environment variables etc.)
- Familiarity with containerization technologies (e.g., Docker).
- Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).
- Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
- Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.
- Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence.
- Strong ability to self-organize and to work independently as part of a remote team Knowledge of version control systems, particularly Git
- Familiarity with programming languages such as Golang or Python
Nice-to-have:
- Experience managing distributed databases or large-scale data storage systems. Knowledge of security best practices in cloud environments.
- Experience with scripting languages like Python or Bash.
- Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. Experience working with GitOps
- Strong programming skills in Golang, with experience in developing automation tools, scripts, or services.
Location: EU Timezone, preferably within the EU itself (Remote)
What Makes Arango Special?
At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.
Working at Arango means:
- Contributing to cutting-edge AI and data infrastructure
- Collaborating with experienced engineers, marketers, and product leaders
- Helping shape how enterprises build AI-powered applications
If you're excited about the intersection of AI, data, and social media, we'd love to hear from you.
$104.9k - $174.7k
Are you passionate about improving reliability, scalability, and resilience in complex database... ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident... ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related...SuggestedFull timeLocal area$176k - $282k
...Well-Architected Framework principles across operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability.Perform Site Reliability Engineering (SRE) functions, including automation of operational tasks, system hardening, and...SuggestedContract workShift work$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours$105.6k - $145.2k
Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience supporting multiple connected Cloud-based products? Trimble is a global technology...SuggestedOngoing contractFull timeWork at officeLocal areaWorldwide$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide$133k - $190k
Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is...Full time$112k - $179k
...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems...Contract workWorldwideShift work$123k - $165k
Job Posting Title:Site Reliability Engineer IIReq ID:10143234Job Description:Department/Group OverviewOur engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational...Full time- ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower businesses... ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely...Temporary workLocal areaWorldwide
$152.6k - $191.5k
...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing... ...improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior individual...Full timeWork at officeDay shift- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...Full time
$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites.About the RoleFivetran is building data...Full timeWork at officeRemote work- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the IP, you will solve complex and broad business problems with simple and straightforward solutions...
$179.2k - $268.8k
...sensors and compute systems, test operations, systems and safety engineering - all dedicated to redefining the relationship between... ...in Dearborn, Mich., and Palo Alto, Calif.Meet the team:As a Site Reliability Engineer on the team, you will be responsible for helping to...Permanent employmentFull timeWork at officeImmediate startVisa sponsorship$102.1k - $202.2k
...per yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole type... ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...workloads. As a Site Reliability Engineer II, you will take ownership of reliability...Ongoing contractWork at officeLocal area- Job ID: 20641226Reference Number: 23-01616Title: Site reliability EngineerLocation: Iselin, NJ, 08830Posted Date: 2023-10-10Contact: Shyam MaramContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN StaffingSRE DevOps lead Lebanon NJ ( Need Local...Work experience placementLocal area
- ...high traffic, business critical internet site communications and/or network-based (... ...teams to ensure software is designed for reliability, scalability, and operational efficiency... ...Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent...Full timeLive outLocal areaFlexible hours
- ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology... ..., 2-days a week in office#LI-KC1#GMFjobsAbout The Role: The Site Reliability Engineer under the general direction from the leadership...Work experience placementH1bWork at officeRemote workVisa sponsorshipFlexible hoursShift work2 days per week
$102.1k - $202.2k
...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewAre... ...no further than the Microsoft Defender engineering team. We are looking for a Site...Ongoing contractLocal area3 days per week$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work$104.43k - $156.65k
...Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual... ..., Fox, Disney, NBC, Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our mission to deliver seamless...Permanent employmentFull timeWork at officeRemote workWorldwideFlexible hours$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...Full timeWork at officeLocal areaRemote workWork from home$62k - $141k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have...Full timeContract workPart timeWork at officeLocal areaRemote work$100k - $125k
...a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?ResponsibilitiesAs a Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS platform...Temporary workCasual workWorldwide- ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States (US).Position OverviewWe are seeking a highly skilled Site...Full timeTemporary workWork at officeRemote workFlexible hours
- ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible... ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry...Flexible hoursWeekend work
$145.7k - $218.5k
...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to... ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated...Work experience placementShift work$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer United States
- lead site reliability engineer United States
- site reliability engineering manager United States
- site reliability engineer remote United States
- site reliability engineer sre United States
- website content developer United States
- after school site coordinator United States
- site leader United States
- site merchandiser United States
- on-site clinical research associate (traveling/remote) United States
