Staff Site Reliability Engineer
$108k - $180kSHEIN
Staff Site Reliability Engineer
San Diego
SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than 15,000 employees operating from offices around the world, SHEIN is committed to making the beauty of fashion accessible to all, promoting its industry-leading, on-demand production methodology, for a smarter, future-ready industry.
Position Summary
We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely. In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations. At the Staff level, you will also provide technical leadership, influencing platform architecture, reliability strategy, and operational standards across the organization.
The SRE team owns and maintains critical open-source and in-house technologies that underpin the platform and serves as a core contributor to major engineering initiatives. We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.
You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling—including metrics, logs, traces, alerting, and automation—providing deep visibility into system behavior. Through hands-on engineering and operational excellence, you will proactively identify risks and failure modes, help prevent incidents before they occur, and lead fast, effective responses when they do. To succeed in this role, you will combine strong software engineering skills, solid to deep expertise in Linux, networking, and distributed systems, and a passion for solving problems of scale, complexity, and reliability. Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.
Job Responsibilities
- Keep SHEIN's mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
- Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
- Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
- Design, build, and maintain observability solutions (metrics, logs, traces, alerting), incorporating AI-powered anomaly detection and intelligent alert correlation to surface actionable signals from high-volume telemetry, improving system visibility and resiliency.
- Automate operational workflows and eliminate manual toil through scripting, tooling, and process improvements, including the use of AI-assisted development tools (e.g., Claude Code) to accelerate the building and iteration of internal operational platforms.
- Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.
- Mentor Senior and mid-level SREs, raising the overall technical bar and operational maturity of the team.
- Lead efforts to modernize the platform in alignment with industry best practices and evolving technology standards.
Job Requirements
- Bachelor's degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- 6+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Experience applying AI/LLM-powered tools to reliability engineering, including designing and building automation or internal tools using AI-assisted development tools (e.g., Claude Code).
- Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
- Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.
- Strong software engineering skills with experience building automation, tooling, or platforms in languages such as Python or Go.
- Experience operating or supporting open-source infrastructure components such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper, etc.
- Experience with observability and monitoring systems (Prometheus, Grafana, Zabbix, etc.) and performance analysis.
- Familiarity with Git, CI/CD pipelines, and configuration management tools (e.g., Ansible).
- A strong sense of ownership, a systematic approach to problem-solving, and a passion for making systems more reliable.
- Strong communication skills and the ability to collaborate effectively with geographically distributed teams.
Nice to Have
- Bilingual fluency in Mandarin and English.
- Kubernetes Administrator certification or equivalent real-world experience.
- Experience operating big data platforms (Hadoop, Yarn, HBase, Hive, Spark).
- Experience applying AI/LLM-powered tools to reliability engineering, including designing and building automation or internal tools using AI-assisted development platforms (e.g., Claude Code).
Benefits and Perks
- Bonus and RSU eligible
- Healthcare (medical, dental, vision, prescription drugs)
- Health Savings Account with Employer Funding
- Flexible Spending Accounts (Healthcare and Dependent care)
- Company-Paid Basic Life/AD&D insurance
- Company-Paid Short-Term and Long-Term Disability
- Voluntary Benefit Offerings (Voluntary Life/AD&D, Hospital Indemnity, Critical Illness, and Accident)
- Employee Assistance Program
- Business Travel Accident Insurance
- 401(k) Savings Plan with discretionary company match and access to a financial advisor
- Vacation, paid holidays, floating holiday and sick days
- Employee discounts
- Free weekly catered lunch
- Dog-friendly office (available at select locations)
- Free gym access (available at select locations)
- Free swag giveaways
- Annual Holiday Party
- Invitations to pop-ups and other company events
- Complimentary daily office snacks and beverages
Pay Range
$108,000 - $180,000 USD
$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...SuggestedFull timeFlexible hours$92.4k - $148.8k
...Senior Site Reliability Engineer San Diego About SHEIN SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than 1...SuggestedTemporary workWork at officeWorldwideFlexible hours- ...contributed to the FaceID and FaceKit project in the past and more recently the new LIDAR iPad sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the next level. In this role, you will help lead our cloud based infrastructure team for...SuggestedWork experience placement
- ...Site Reliability Engineer This position will primarily focus on providing design and implementation expertise on infrastructure provisioning, management and lifecycle implementation of cloud components and services, containers and other critical concepts of DevSecOps...SuggestedShift work
- ...We are seeking a Senior Site Reliability Engineer (SRE) to join a high-impact Platform Engineering team focused on building scalable cloud infrastructure, reusable platform capabilities, and automation frameworks that enable software engineers to develop and deploy applications...Suggested
- ...with preference to candidates located in San Diego, CA, Norfolk, VA or Charleston, SC Position Overview: The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government...Contract workRemote work
$87.1k - $157.45k
More About the Role:Leidos is seeking a Site Reliability Engineer (SRE) Data Engineer supporting the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the Leidos team delivers the core backbone of the Navy...Full time$140k - $190k
...expert in version control, continuous integration, and build systems who is equally comfortable driving a release end to end and engineering the systems that streamline it. You will work closely with our customers, from the teams whose components we release to the downstream...Full timeTemporary workRelocation package- ...independently, with some supervision. Participates in design reviews and project meetings. Telecommuting permitted. Will accept a Bachelor's Degree (or foreign academic equivalent) in Electrical Engineering, Computer Engineering, Computer Science or related degree field.Full timeRemote work
- ...Operational Architecture (NOA), working alongside engineers and operators to deliver resilient,... ...and analytics to ensure accuracy, reliability, and robustness. Stay updated with the... ...preferred Location/Address: ~100% On-site ~ San Diego, CA, NAVWAR ~ Arlington,...Full time
- ...hardworking and task-oriented. Don't Wait! Fill out a Profile Now! MyJobResource is a staffing and recruitment industry job search engine. We specialize in finding the exact company to suit your needs. We help match job seekers to the right jobs in either full-time or...Full timeTemporary workPart time
$146.3k - $219.5k
...specifically the Pricing and Estimating group, where we partner with engineers, manufacturing, supply chain, program management, and others to... ...who do it well shape decisions that play out over years.As a Staff Proposal Analyst, you'll join a business unit within the Air...Full timeRelocation packageShift work$50k
...Design and conduct training sessions and workshops for Legal, Claims, and related departments. Mentor junior attorneys and legal support staff, offering guidance and direction to elevate team performance.Documentation and Compliance: Ensure all legal documents are...Local areaRemote workFlexible hours$215k - $260k
...impact operations.Your work will directly safeguard the safety, reliability, and scaling of the Zoox autonomous fleet. Robust, low-... ...Zoox’s connectivity systemsQualificationsM.S. in Electrical Engineering, Telecommunications, or a similar discipline with 8+ years of...Full timeTemporary workRelocation package$111.3k - $166.9k
...Company: Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group Software Engineering General Summary:... ...To all Staffing and Recruiting Agencies : Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and...Full timeWork experience placementWork from home- ...and software power/thermal management techniques Apply deep expertise in power, performance and thermal domain to achieve critical engineering goals. Lead the ideation, prototyping, and development of key power, performance, and thermal features, ensuring product KPIs are...Full time
$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting – Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Full timeWork at office$111.3k - $166.9k
....\n\n## Job Area:\n\nEngineering Group, Engineering Group \u003e Software Engineering\n\nGeneral... ...sponsorship.\n\nDrive innovation and reliability for Snapdragon SoC-based platforms at... ...Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at...Full timeWork experience placementWork from home- ...through a hybrid approach. Teradata delivers real business value with AI.What You Will Do We are seeking a Senior Platform Systems Engineer to help architect and evolve next-generation AI and database infrastructure platforms used in enterprise and hyperscale datacenter...
$208k - $300k
Foster City, CA / San Diego, CASoftware - Software Systems /Full-time /HybridSimulation is central to ensuring the safety and scalability of autonomous vehicles at Zoox. In this role, you will develop metrics and evaluation pipelines to assess simulator fidelity and guide...Full timeTemporary workRelocation package$106.8k - $133.5k
...make recommendations.Advanced knowledge of plant and equipment reliability.Travel between facilities to support all functional areas.... ...equivalent experience required.Minimum of 7-10+ years of relevant engineering experience including a minimum of 4 years cGMP experience....Full timeWork at officeFlexible hours$125k - $175k
...Join our fast-paced and passionate team as a Senior Software Engineer. As we scale, you will be instrumental in building our foundation... ...for commercial and internal applications. Design and develop reliable software components that interact with hardware devices....Full timeLocal areaFlexible hours$115k - $185k
...you’re never going alone. Because there’s too much at stake to go solo. Our Radio Products Team is seeking a hybrid Software Engineer, Modem BSP. You would be responsible for working on next generation self-networking hand-held radios for domestic and foreign defense...Permanent employmentFull timeWork experience placementWork at officeWorldwide- Company Description AG Technologies was founded as a software solutions company in 2008 & has its corporate headquarters at Chesterfield, Missouri with branches within the US and India. Over the years the organization has expanded into various market segments & activities...Full time
- ...GA with prime emphasis on the following service offerings: Staff Augmentation Lifecycle IT solutions Application Development... ...Test Automation Job Description Job Position: Software Engineer (DevOps) Location: San Diego, CA Duration: 6 Months...Full time
$150k - $190k
...funding for life-saving programs like In-Home Supportive Services (IHSS) and subsidized family childcare. Job Information: Job Title: Staff Counsel Job Type: Exempt (Salary) Department: Executive Director Reports to: Executive Director Schedule: Monday to Friday, 9:00 AM...Contract workWork at officeLocal areaMonday to Friday$140k - $175k
...critical assets to lead in the race for technological and operational superiority from ground to space. We are looking for a SW Engineer to join our team of designers of cutting-edge components for space and national security applications, including Software-Defined...Long term contractPermanent employmentFull timeContract workFlexible hours- ...on laptop/desktop, tablets, phones and devices yet to be thought of Continue to improve upon your already excellent software engineering knowledge Create the best software of your life Work in a Behavior/Test Driven Development environment Be excited by the...Full time
$135k - $156k
...San Diego, California . Summary Designs, develops, and implements enterprise-grade software solutions supporting FRCSW engineering, logistics, and business systems. Leads full software lifecycle task requirements analysis, architecture, coding, testing, integration...Full timeFor contractorsFor subcontractor- ...reshaping it. Our dedication to merging the prowess of humans and machines to solve complex problems has set us apart in designing and engineering solutions for the Department of Defense (DoD) networks. Here, every challenge is an opportunity to advance, and every solution is...Full timeFor contractorsLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- staff engineer San Diego, CA
- assistant engineer San Diego, CA
- engineering aide San Diego, CA
- senior staff engineer San Diego, CA
- senior staff systems engineer San Diego, CA
- assistant electrical engineer San Diego, CA
- software engineer staff San Diego, CA
- technology administrator San Diego, CA
- project engineer assistant project manager San Diego, CA
- junior website developer San Diego, CA


