SRE
InterSources Inc
SRE Production Support Engineer
Location: Plano, TX
Duration: 6 Months (Contract to hire)
Interview Process:
1st round - Zoom
2nd round – In Person
Role Overview:
Position is part of the Central Site Reliability Engineering (SRE) Team.
Looking for a strong individual contributor, not someone who simply follows instructions.
Candidate should have a Software Engineering background, with a strong focus on production environments and application reliability.
This is not a traditional support or operations role.
Required Skills:
Solid understanding of Site Reliability Engineering (SRE) principles.
Experience managing and maintaining production applications ("care and feeding" of applications).
Ability to review and manage production backlogs.
Production-focused mindset with experience in reliability, stability, and operational excellence.
Understanding of how applications are deployed, monitored, and maintained.
Strong proficiency in Java coding and scripting – Python, Shell Scripting
Good understanding of DevOps concepts and practices.
Familiarity with .NET applications is beneficial.
Experience with any major cloud platform is acceptable, including:
AWS
Azure
Google Cloud Platform (GCP)
IBM Cloud
Dell Technologies Cloud
The specific cloud platform is less important than understanding how applications are deployed, managed, and operated in a cloud environment.
Role Expectations:
This is not a shift-based position.
The role does not involve continuously monitoring dashboards or providing traditional production support.
The focus is on improving application reliability through engineering and automation.
Candidate does not need to know the exact source code of every application but should understand how applications function in production and how to improve their reliability.
Job Description:
Expert in at least one technology and design technique as well as experience working across large environments with multiple operating systems/infrastructure for large-scale programs (e.g., Expert Engineers) starting to be firm-wide resources working on projects across Client
Is multi-skilled with expertise across software development lifecycle and toolset
May be recognized as a leader in Agile and cultivating teams working in Agile frameworks
Sought out as coach for at least one technical skill
Strong understanding of techniques such as Continuous Integration, Continuous Delivery, Test Driven Development, Cloud Development, resiliency, security
Stays abreast of cutting-edge technologies/trends and uses experience to influence application of those technologies/trends to support the business; may give speeches and outside the firm, writes articles
Roles and Responsibilities:
Executes standard software solutions, design, development, and technical troubleshooting
Writes secure and high-quality code using the syntax of at least one programming language with limited guidance
Designs, develops, codes, and troubleshoots with consideration of upstream and downstream systems and technical implications
Applies knowledge of tools within the Software Development Life Cycle toolchain to improve the value realized by automation
Applies technical troubleshooting to break down solutions and solve technical problems of basic complexity
Gathers, analyzes, and draws conclusions from large, diverse data sets to identify problems and contribute to decision-making in service of secure, stable application development
Learns and applies system processes, methodologies, and skills for the development of secure, stable code and systems
Adds to team culture of diversity, equity, inclusion, and respect
Additional Skills:
Formal training or certification in software engineering /Site Reliability Engineering concepts and 6 plus years of applied experience.
Hands-on practical experience in system design, application development, testing, and operational stability
Exposure to product engineering or production/Platform support activities with a good understanding on scalability, security, and reliability.
Experience in developing, debugging, and maintaining code in a large corporate environment with one or more modern programming languages and database querying languages
Demonstrable ability to code in one or more languages
Experience across the whole Software Development Life Cycle
Exposure to agile methodologies such as CI/CD, Application Resiliency, and Security
Emerging knowledge of software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, mobile, etc.)
About Us:
InterSources Inc, a Certified Diverse Supplier, was founded in 2007 and offers innovative solutions to help clients with Digital Transformations across various domains and industries. Our history spans over 16 years and today we are an Award-Winning Global Software Consultancy solving complex problems with technology. We recognize that our employees and our clients are our strengths as the diverse talents and opportunities they bring to the table enable us to grow as a global platform and they are causally linked with our success. We provide strategic and technical advice, and we have expertise in areas covering Artificial Intelligence, Cloud Migration, Custom Software Development, Data Analytics Infrastructure & Cloud Solutions, Cyber Security Services, etc. We make reasonable accommodations for clients and employees and we do not discriminate based on any protected attribute including race, religion, color, national origin, gender sexual orientation, gender identity, age, or marital status. We also are a Google Cloud partner company. We align strategy with execution and provide secure service solutions by developing and using the latest technologies that thrive our resources to deliver industry-leading capabilities to our clients and customers, making it convenient for our clients to do business with InterSources Inc. Our teams also drive growth by refining technology-driven client experiences that put the users first, providing an unparalleled experience. This results in strengthening the core technologies of clients, enabling them to scale with flexibility, create seamless digital experiences and build lifelong relationships.
- Job Title SRE Lead Location Plano, TX - Onsite Job Description We are seeking an experienced 13 to 18 years of experience to join our team. The ideal candidate will have expertise in AWS, SRE, and Datadog, and a background in the automotive industry is a plus. This hybrid...SuggestedDay shift
- We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.As a Software Engineer III at JPMorgan Chase, you serve as a seasoned member of an agile team to design and deliver trusted market-leading technology products...Suggested
- ...Engineering (SPE) team within the Toyota Financial Services CTO PM organization is seeking a Senior Product Manager to own and transform the SRE & Observability product area with an expanded charter to elevate the space and shape broader AIOps strategy for TFS Team Members. You...SuggestedH1bShift work
- Toyota North America is seeking a Senior Product Manager to own and transform the SRE & Observability space, with an expanded charter to shape broader AIOps strategy for TFS team members. You will lead end-to-end product vision across production and sub-productions, translating...Suggested
- Toyota Financial Services seeks a Senior Product Manager to own the SRE & Observability space, driving a product-driven model with reliability outcomes and business impact. You will lead roadmap prioritization, VOC synthesis, acceptance criteria, and metrics across SRE,...Suggested
$155.4k - $261.1k
...communication skills to explain complex issues clearly Required: ~7+ years in Systems Engineering, ITSM, RM/CM ~ Background in SRE, Support or QA ~ One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana ~ Hands-on experience and...Permanent employmentFull timeTemporary workWork at officeLocal areaRelocationShift work$197.3k - $225.1k
...Overview Lead Software Engineer, Full Stack (Java, Python, SRE, AWS, AI) (Cloud Operations Resilience Engineering) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive...Full timePart timeInternshipH1bLocal area- Mandatory skills: Azure DevOps (ADO), GitHub & GitHub Actions, JFrog Artifactory Site Reliability Engineer (SRE) - Will require end2end testing, performance testing, failover. - Overall understanding of the Infra and Application to conduct load testing, failure...
- ...Sre Position 8+ years of professional experience processing a culture of learning through the development and sharing of skills, knowledge, process and tools. A driving passion for finding solutions to hard problems at scale and operationalizing them. Exceptional critical...
$66.73 per hour
...Description We are seeking a Site Reliability Engineer (SRE) to help establish and scale our client's Google Cloud Platform (GCP) SRE practice. This is a unique opportunity to join a team during its formative stage and play a key role in building the operational foundation...Contract workTemporary work$87.5k - $125k
Observability Engineer DISH is transforming the future of connectivity. We're doing it by building the country's first virtualized, standalone 5G wireless network from scratch. The foundation of a connected world, it's a network free of the limitations of the past, ...Flexible hoursNight shift- ...delivering measurable improvementsManage feedback and field questions from all levels of stakeholders, within the (Tech, Product, QA, SRE, Design, Arch) tower and across CTO towers to better standardize processes (backlog health, estimation), spec-driven development, and...Temporary workWork at office
- ...Dynatrace with ServiceNow Splunk PagerDuty Jira or similar platformsHandson scripting experience with Python Shell Script or PowerShellFamiliarity with DevOps CICD Infrastructure as Code and Site Reliability Engineering SRE practicesITIL Foundation certification preferred
- ...knowledge of Linux administrationExperience in production support troubleshooting and incident managementGood to HaveExperience with SRE practices observability SLAs reliability engineeringExposure to monitoring tools Prometheus Grafana etcKnowledge of data replication...
- ...root cause analysis, identifying systemic risks, and implementing preventative measures to reduce production incidents.Familiarity with SRE practices such as service level indicators, service level objectives, error budgets, toil reduction, blameless post-incident reviews,...
- ...flexibility for diverse business needs.Serve as the primary observability subject matter expert and trusted advisor to engineering, product, SRE, Cloud Operations, FinOps, governance, and executive leadership teams.Evaluate, recommend, and guide the implementation of...Shift work
- ...operational documentation, runbooks, and knowledge base articles.Collaboration and Stakeholder EngagementPartner with application development, SRE, architecture, security, and platform teams to deliver stable, scalable solutions for our clients.Advise on schema design, data...Permanent employmentFull timeMonday to FridayFlexible hours
$61.37 per hour
DescriptionWe are seeking a Technical Product Owner to partner closely with Engineering and Site Reliability Engineering (SRE) teams in driving infrastructure and cloud platform initiatives. This is a highly technical, engineering-facing role focused on translating business...Contract workTemporary work- ...environments. Oversee cloud engineering, platform engineering, and infrastructure modernization initiatives. Implement modern DevOps, SRE, automation, and continuous delivery practices. Establish engineering standards, governance, and reusable frameworks that improve...Full time
- ...query limits, N+1 mitigation (DataLoader).Proven delivery of API/schema governance, versioning/deprecation, and CI policy gates.Strong SRE practices: SLIs/SLOs, error budgets, OpenTelemetry, data-driven post-incident improvements.Developer productivity: time-to-first-...Contract work
- ...identifying emerging talent, fostering leadership skills, and managing stakeholders.Overview:Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead...Full timeWork at officeDay shift
$117k - $209.33k
...Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique opportunity to help shape how Autodesk deploys, runs, and improves...Full timeFor contractorsRemote work- ...area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.Job responsibilitiesLead SRE practices that balance delivery speed, efficiency, and system stability Partner with engineering peers and senior stakeholders to drive...Work at office
- ...an application or platformDemonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity...Work experience placement
$152.6k - $191.5k
...operationsProvide technical design input for new Azure services and workloads to ensure operational readiness before production adoptionMentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production supportCreate...Full timeWork at officeDay shift- ...of Experience: 8 to 15 yearsSkillsKubernetes and DockerJob SummaryWe are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have extensive experience in maintaining production systems, particularly in cloud...Remote work
- ...certification on site reliability engineering concepts and 3+ years applied experienceExposure to or hands-on experience in supporting SRE practices for Data management/migration platforms and products, with familiarity in on-prem/public cloud infrastructure components...
- ...of Linux administration. Experience in production support, troubleshooting, and incident management. Good to Have Experience with SRE practices - observability, SLAs, reliability engineering. Exposure to monitoring tools such as Prometheus, Grafana, etc. Knowledge...Temporary workLocal area
- ...Infrastructure as Code). Python scripting. AWS, Azure, or GCP. Linux administration. Production support and incident management. Nice to Have SRE practices Prometheus Grafana Data replication Sharding Distributed architectures #J-18808-Ljbffr First Soft Solutions LLCFull time
- ...(schema evolution, backfills, error handling, contract-driven interoperability).Establish production-grade operations and controls (SRE practices, monitoring/on-call, incident response/RCA, auditability, least-privilege, disciplined change management) and deliver governed...Contract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


