Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Reliability Engineer

Software Technology Inc

Splunk Performance EngineerFULLY REMOTE BUT EST HOURS DONT SEND WEST COAST CANDIDATESSTRONG SPLUNKDuties include:Develop and maintain comprehensive monitoring solutions for cloud-based services and applications.Configure monitoring tools and systems to collect relevant metrics, logs, and traces.Create custom monitoring dashboards and reports using Splunk, DataDog, DynaTrace or other tools, to provide real-time insights into system performance and health.Continuously monitor the cloud infrastructure's performance and capacity, anticipating and addressing potential scalability issues.Proactively suggest and implement improvements to enhance the system's reliability, resilience, and fault tolerance.Work on automating tasks to streamline operational processes and reduce manual intervention.Collaborate with cross-functional teams to investigate and resolve critical incidents, ensuring minimal impact on end-users.Work with Problem Management team to complete post-mortem analysis of incidents to identify root causes and implement preventive measures.Understand the overall architecture of our systems to identify gaps in monitoring and troubleshoot issues.Configure and maintain custom dashboards and alerts in various monitoring tools.Create custom reports, deliver report presentations to various stakeholders.Develop scripts for monitoring PowerShell, Python, Shell scripting.Develop metrics for both the business and technical teams to determine the health of systems.Provide on-call support as needed.Leads and coordinates performance engineering for medium to large initiatives.Collect and document expected system performance and operational characteristics.Collect and/or prepare test data for test execution.Develop and execute performance tests including load, stress, endurance, fail-over and interoperability.Conduct technical analysis of performance test results and production systems, and provide recommendations on performance tuning, systems, and infrastructure. Identify, report, and review defects in assessing system performance and stability.Defining the strategy for enabling performance diagnostics and monitoring using an Application Performance Management (APM) tool, other monitoring tools, and diagnostic techniques.Collaborating with developers to promote the concept of performance engineering during all phases of the SDLC to detect and correct performance issues earlier in the lifecycle.Leads peer reviews to ensure the completeness of all test assets created.Resolve performance and stability issues in performance test environment.Develop performance engineering work plan structure and project schedule.Review architectural design for performance risks and potential issues.Prepare capacity analysis when applicable.Minimum Requirements:Requires a BA/BS degree in Information Technology, Computer Science or related field of study and a minimum of 7 years performance engineering and performance testing experience; or any combination of education and experience, which would provide an equivalent background.Preferred Skills, Capabilities and Experiences:Experience managing performance engineering efforts for an application strongly preferred.Proficiency with the following tools is preferred (Splunk, DataDog, DynaTrace among others).Experience managing performance engineering efforts for an application strongly preferred.Knowledge of developing scripts for monitoring (PowerShell, Python and Shell scripting).5 years’ of Splunk programming proficiency is highly preferred.5-6 years’ experience using.NET and Java application and Application Monitoring Tools like App Dynamics or Datadog are highly preferred.Proficiency is performance tuning is preferred.Good understanding of the UI, Middleware and backend Databases

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Reliability Engineer in Washington DC vacancy
  •  ...established and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded in...  ...our team.Barbaricum is seeking an experienced Senior Site Reliability Engineer to support the reliability, availability, automation... 
    Suggested
    Contract work
    For contractors

    Barbaricum

    Washington DC
    3 days ago
  •  ...Mid-Level Reliability EngineerDecision Technologies was founded to help bridge the gap between Department of Defense and Federal Government...  ..., our areas of expertise have expanded to include Systems Engineering, Program Management, In-Service Engineering, Equipment Repair... 
    Suggested
    Contract work
    For subcontractor
    Work at office
    Remote work

    Decision Technologies, Inc.

    Arlington, VA
    2 days ago
  •  ...Management, Compliance, Business Process, IT Effectiveness, Engineering, Environmental, Sustainability, and Human Capital. We help forward...  ...at DescriptionProSidian Seeks a Downstream Oil & Gas Reliability Engineer | Technical Due Diligence & Engineering Validation For... 
    Suggested
    Full time
    Contract work
    Temporary work
    For contractors
    Work at office
    Remote work
    Flexible hours

    Prosidian Consultng

    Washington DC
    4 days ago
  • $96k - $140k

     ...partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleProduct Reliability Engineers (PREs) are responsible for the health, performance, and stability of the services that power services at Palantir. PREs take... 
    Suggested
    Permanent employment
    Full time
    Fixed term contract
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package
    Shift work

    Palantir Technologies

    Washington DC
    1 day ago
  • $108k - $162k

    DescriptionWe're seeking a Regional Reliability Engineering Manager to help drive reliability excellence across lumber manufacturing operations in the Pacific Northwest and Alberta, Canada region. In this highly visible role, you'll partner with mill leaders, reliability... 
    Suggested
    Temporary work
    Remote work
    Work from home

    Weyerhaeuser

    Washington DC
    1 day ago
  • $165k - $230k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts. While... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    3 days ago
  • $185k - $230k

    As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver technical solutions within the highest levels of the federal government.We know that you can’t have great technology services... 
    Full time
    Local area
    Immediate start

    MetroStar Systems

    Washington DC
    1 day ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    1 day ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions...  ...missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    3 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    2 days ago
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    3 days ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS platform... 
    Temporary work

    Kong

    Washington DC
    6 hours ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    6 hours ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  • $125k - $185k

     ...drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  • $150k - $180k

     ...redefining what’s possible in remote sensing, you belong here at Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems. In... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    6 hours ago
  • $160k - $220k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE)  Experience Level: Mid–Senior (4+ years PostgreSQL experience) About the Role We are looking for a highly... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    more than 2 months ago
  • $166k - $220k

     ...or failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services & infrastructure...  ...& Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril’s production... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    3 days ago
  • $207k - $284.9k

     ...this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    3 days ago
  • $174k - $238k

     .... We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    6 hours ago
  •  ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats... 

    Cyrad Solutions LLC

    Washington DC
    1 day ago
  • $100k - $110k

     ...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will... 
    Permanent employment
    Remote work
    Flexible hours

    GrabJobs

    Washington DC
    2 days ago
  • $165k - $230k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy... 
    Permanent employment
    Full time
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Washington DC
    2 days ago
  • $90k - $105k

     ...digital platform for light—and redefining what’s possible in the optical age. Job Description: We are looking for a reliability engineer to own the reliability of the LCM itself — the chip at the core of everything Lumotive ships. The LCM is a new class of semiconductor... 
    Full time
    Shift work

    Lumotive

    Washington DC
    1 day ago
  • $82.3k

     ...inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer – Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud... 
    Temporary work
    Work at office
    Remote work
    Worldwide
    3 days per week

    GrabJobs

    Washington DC
    2 days ago
  •  ...certificates. Qualifications & Requirements ~ Bachelor’s degree in Computer Science, Engineering, or a related technical field. ~3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or infrastructure-focused roles. ~ Hands-on... 
    Temporary work

    2T Consulting

    Washington DC
    3 days ago
  • $145k - $175k

     ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data... 
    Full time
    Remote work

    GrabJobs

    Washington DC
    2 days ago
  • $175k - $195k

     ...they love Filevine products need key features. They need to be reliable, scalable, performant, cost effective, secure and they need to...  ...team is responsible for thinking through these problems and engineering solutions to them. We hire excellent engineers who apply software... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    Filevine

    Washington DC
    2 days ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    25 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!