Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Kaav Inc.

Job Description:
  • DevOps
  • Docker/Kubernetes
  • Splunk/Dynatrace
  • Ansible/GitHub
  • Java Full Stack Applications Deployment
  • Responsible for reliability and support of Container Platform on-prem and external clouds (Azure /AWS /Google)
  • Monitor and troubleshoot Container platform environment performance issues, connectivity issues, security issues, etc.
  • Perform deep dives into systemic and latent reliability issues, Incident management, problem management
  • Identifying, analyzing, and resolving infrastructure vulnerabilities and application deployment issues.
  • Perform blameless RCA, partner with engineering and operation teams across the organization to roll out fixes.
  • Responsible for application onboarding and provide troubleshooting support through the lifecycle of the applications on the container platform.
  • Identify and drive opportunities to improve automation to reduce TOIL and improve operational excellence.
  • Partner with risk, and compliance teams to bring visibility and implement right controls and remediation of vulnerabilities.
  • Ensure resiliency during implementation and identify/fix resiliency problems by collaborating with engineering teams.
  • Be a key stakeholder in the design of cloud services and work with Architecture, engineering, product teams
  • Participate in 24x7 on-call coverage follow the sun model
  • S /MS degree in Computer Science or related technical field involving systems or equivalent practical experience.
  • Minimum 5+ years of hands-on experience supporting Kubernetes /Openshift / RKE / EKS Container platform.
  • Experience with Python, Ansible, Golang, and shell scripting
  • Experience with Splunk/Dynatrace
  • Experience with Java Full Stack Applications Deployment
  • Kubernetes /Openshift /Terraform certifications are a plus
  • Strong experience in major services related to Compute, Storage, Network and Security
  • Experience with monitoring tools like Prometheus and Dynatrace, as well as cloud native tools like Azure Monitor and Log Analytics
  • Strong understanding and background of working with a complex IAM infrastructure, including Active Directory, Azure AD Connect, Azure AD, and Ping Identity or other SSO solutions.
  • Advanced knowledge of Linux OS, DNS, DHCP, Kerberos and Windows Authentication
  • Experience with CI/CD tools git /Jenkins, GitOps model
  • Excellent understanding of Linux /Windows operating systems administration
  • Experience in Container security and vulnerability remediation.
  • Experience with Ansible/GitHub
  • Systematic problem-solving approach, sense of ownership and drive
  • Ability to juggle competing priorities and adapt to changes in project scope.
  • Excellent interpersonal, organizational and communication (written, verbal, and presentation) skills are a must.
  • Proven ability to work independently with minimal supervision and as part of a team with direct responsibilities.
  • Experience in Openshift, RKE, CSP Kubernetes services such as AKS and EKS
  • Experience in Terraform, ArgoCD, Tekton, and K-native technologies.
  • Experience in agile deployment methodologies (GitOps)
  • Knowledge of various container runtimes
  • Familiarity with the operator deployment pattern.
  • Experience working in a highly available multi-datacenter environment
  • Experience working with monitoring tools such as Prometheus, Splunk, Dynatrace, Sysdig, or similar tools.
  • Understanding of cost management, inventory management, FinOps model
Required Skills : Grafana,Docker,Kubernetes,ansible,Java
Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis is a high PRIORITY requisition. This is a PROACTIVE requisition
Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  •  ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Suggested
    Shift work

    J.P. Morgan

    New York, NY
    5 days ago
  •  ...collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology, liquidity risk team, you are the non-functional requirement owner... 
    Suggested

    J.P. Morgan

    New York, NY
    2 days ago
  •  ...and shape the future of technology at a globally recognized firm, driven by pride in ownership. As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets team, you are the non-functional requirement owner and... 
    Suggested
    Bank staff
    Shift work

    J.P. Morgan

    New York, NY
    17 days ago
  •  ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You... 
    Suggested
    Full time
    Work at office

    Dune Security

    New York, NY
    16 hours ago
  •  ...millions of people around the world. About the role Novellia is a Series A health tech startup, and we're hiring our first Site Reliability Engineer. You'll join Platform Engineering as its second member, working directly with the Head of Platform Engineering to... 
    Suggested
    Flexible hours

    Novellia

    New York, NY
    1 day ago
  • $100k - $250k

     ...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial...  ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics... 
    Local area

    Kalshi

    New York, NY
    1 day ago
  • $164k - $205k

     ...ensuring high availability and performance Design intelligent alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on operational concerns Automate incident response workflows and build... 
    Work experience placement
    Summer holiday
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    BetterUp

    New York, NY
    1 day ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around... 
    Local area
    Flexible hours

    MongoDB

    New York, NY
    1 day ago
  •  ...configure the monitoring and alerting metrics so the support engineers can proactively and timely validate, troubleshoot and...  ...monitoring high availability critical application compon ents.1+ Years in Site Reliability Engineering organization prefe #J-18808-Ljbffr... 
    Work experience placement

    PineQ Lab Technology

    New York, NY
    1 day ago
  • $176.75k - $209.1k

     ...Site Reliability Engineer At Peloton, we view Platform as a Product. A phenomenal platform unlocks speed of development and learning. It allows us to scale easily, enabling our engineers to maximize attention on new features and capabilities. A key to crafting a phenomenal... 
    Temporary work

    Peloton

    New York, NY
    3 days ago
  • $150k - $175k

     ...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed... 
    Remote work

    ASAPP

    New York, NY
    2 days ago
  •  ...A leading quantitative trading firm is hiring a Site Reliability Engineer to help build and operate the critical systems powering its global trading infrastructure. This is a high-impact role focused on reliability, automation, scalability, and performance across mission... 

    Goliath Partners LP

    New York, NY
    1 day ago
  • $123k - $165k

     ...Site Reliability Engineer II Our engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational support to backend service development teams. Technology is... 

    Disney

    New York, NY
    3 days ago
  •  ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews...  ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience... 

    Bahwan CyberTek

    New York, NY
    4 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient... 

    TechChain Talent

    New York, NY
    1 day ago
  •  ...New York City or Chicago (Hybrid) A technology-driven investment firm is expanding its Platform Engineering organization and is seeking an experienced Senior Site Reliability Engineer to help shape reliability practices across its infrastructure and production... 

    Mission Staffing

    New York, NY
    1 day ago
  •  ...Responsibilities Improve the reliability of mission-critical solutions, applications, and platforms Software development for enterprises...  ...Windows and Linux Years of Experience: 5 Years of Software Engineering Seniority level Mid-Senior level Employment type Full-time Job... 
    Full time
    Work experience placement

    InterEx Group

    New York, NY
    1 day ago
  •  ...Site Reliability Engineer Discover your future at Citi. Working at Citi is far more than just a job. A career with us means joining a team of more than 230,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to... 

    Citi

    New York, NY
    4 days ago
  • $89k - $178k

     ...industry. Learn more at What You’ll Do Build and maintain the reliability, scalability, and performance of our digital media...  ...prevent recurrence Required Experience & Skills 4+ years in Site Reliability Engineering, DevOps, or related operational roles with proven... 

    DoubleVerify

    New York, NY
    1 day ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native software applications that serve millions of customers and process billions of dollars in payments. We’re looking for a seasoned... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    ZIP

    New York, NY
    4 days ago
  • $400k

     ...Salary : Up to $400,000 Total Compensation Senior Site Reliability Engineer We are working with a leading trading technology firm building a high-performance infrastructure engineering team focused on reliability, automation, and large-scale platform resilience. This is... 

    Hamilton Barnes ?

    New York, NY
    1 day ago
  • $120k - $160k

     ...and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production operations and software development as you improve, manage, and monitor... 
    Work at office
    Local area

    The Voleon Group

    New York, NY
    1 day ago
  • $150k - $200k

     ...Site Reliability Engineer at Triomics (W21) $150K - $200K AI Agents for Oncology EHRs Triomics is building the agentic AI layer for oncology electronic health records (EHRs). Cancer hospitals spend billions on highly trained staff manually reading unstructured patient... 
    Full time
    Work at office
    Remote work
    Day shift

    Triomics

    New York, NY
    4 days ago
  •  ...Our client is looking for an Infrastructure Engineer to build and operate the foundational systems that power our data, analytics, and...  ...container orchestration, CI/CD, and monitoring keeping our platform reliable, scalable, and secure. If you are excited about AI and want to... 
    Flexible hours
    3 days per week

    The Phoenix Group

    New York, NY
    1 day ago
  •  ...firm that has been operating at scale since 2014, with millions of global users and a reputation for rigorous engineering. The Role As Senior Site Reliability Engineer , you will own the infrastructure foundation that the entire engineering organization depends on. This... 

    Harrison Clarke

    New York, NY
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE / Infrastructure) Role Overview We’re hiring a Senior SRE to build and scale the infrastructure behind a high-growth, production system. You’ll ensure reliability, performance, and scalability as the platform grows from early traction... 

    The Cypress Group

    New York, NY
    1 day ago
  •  ...A leading quantitative trading firm is seeking a Senior Site Reliability Engineer to build and evolve the reliability, observability, and automation capabilities powering a highly performance-sensitive trading environment. Working at the intersection of software and infrastructure... 

    Acquire Me

    New York, NY
    1 day ago
  •  ...significantly reduces costs and improves the critically important 24x7 performance for building owners, developers and tenants. Site Reliability Engineer II The SRE II sits at the intersection of software engineering and platform operations. You will own the reliability,... 
    Remote work

    Kastle Systems

    New York, NY
    1 day ago
  •  ...future of legal tech — we’re defining it. Ready to join us in building the intelligent future of law? The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering hub, sitting within Foundations. You'll own critical services... 
    Work at office

    Legora AB

    New York, NY
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling... 
    Flexible hours

    Baseten

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!