Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer with ML platform - Only W2

Saransh Inc

Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud solutions, build MLOps on cloud AWS Data science model containerization, deployment using docker, VLLM, Kubernetes Communicate with a team of data scientists, data engineers and architects, document the processes Develop and deploy scalable tools and services for our clients to handle machine learning training and inference Knowledge of ML models and LLM Qualifications 6+ years of experience in ML Ops with strong knowledge in Kubernetes, Python, MongoDB and AWS Good understanding of Apache SOLR Proficient with Linux administration Knowledge of ML models and LLM Ability to understand tools used by data scientists and experience with software development and test automation Ability to design and implement cloud solutions and ability to build MLOps pipelines on cloud solutions (AWS) Experience working with cloud computing and database systems Experience building custom integrations between cloud-based systems using APIs Experience developing and maintaining ML systems built with open-source tools Experience with MLOps Frameworks like Kubeflow, MLFlow, DataRobot, Airflow etc., experience with Docker and Kubernetes Experience developing containers and Kubernetes in cloud computing environments Familiarity with one or more data-oriented workflow orchestration frameworks (Kubeflow, Airflow, Argo, etc.) Ability to translate business needs to technical requirements Strong understanding of software testing, benchmarking, and continuous integration Exposure to machine learning methodology and best practices Good communication skills and ability to work in a team #J-18808-Ljbffr Saransh Inc

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer with ML platform - Only W2 in Sunnyvale, CA vacancy
  • $152k - $241.5k

     ...of our global services platform. At NVIDIA, you’ll keep...  ...they integrate cleanly with HPC schedulers, storage...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...driven operations (AIOps/ML-driven signals) that...  ...or Ruby.Mentored other engineers and influenced technical... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

     ...Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara...  ...-board new applications, AI/ML services, and model...  ...site production environment with a real passion for CDN automation...  ...knowledge of the Kubernetes Platform, deployments, and cloud-native... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  •  ...Description Job Description Site Reliability Engineer Onsite- Bay Area, CA...  ...reliability, and uptime across platforms. Handle infrastructure...  ...Engineering ~ Solid experience with GCP or AWS (hybrid/on-prem a...  ...Experience with scalable GPU infrastructure for AI/ML... 
    Platform

    Amiri Recruiting

    Mountain View, CA
    more than 2 months ago
  • $169k - $338k

     ...Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization, you will...  ...cutting-edge machine learning platforms and autonomous agents that revolutionize...  ...organization is built with hybrid systems and software engineers... 
    Platform
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    4 days ago
  •  ...is currently Tuesday.Engineering at Lambda is responsible...  ...customers with Kubernetes questions,...  ...automate the validation of platform quality.Design, build,...  ...workloads, and platform reliability.You6+ years of experience...  ...experienceExposure to HPC clusters, AI/ML workloads, or large-... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $122.5k - $175k

     ...-native Zero Trust Exchange platform. This innovation protects our...  ...make an impact quickly and with high quality. To do this, we...  ...RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a...  ...understanding of AI/ML technologies and experience... 
    Platform
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    3 days ago
  • Elevate your engineering prowess to unprecedented levels...  ...among the top echelon in site reliability. As a Senior Lead Site...  ...the Infrastructure Platforms and Foundational Services...  ...(IPFS) team, you work with your fellow...  ...Generation)Familiarity with AI/ML model building,... 
    Platform

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure,...  ...organization within NVIDIA. This group works with various other groups within NVIDIA...  ..., and Android. It supports hardware platforms including NVIDIA GPUs and Tegra... 
    Platform
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $118.66k - $259.2k

     ...Site Reliability Engineer - AML Global Recommendation - USDS About the Team: Site...  ...a massively distributed AI/ML recommendation system for...  ...systems. Collaborate closely with software engineering teams...  ...and protection of the TikTok platform and U.S. user data, so... 
    Platform
    Temporary work
    Work at office
    Shift work
    3 days per week

    TikTok

    San Jose, CA
    5 hours ago
  • $230k - $250k

     ...autonomous networking, giving engineers and AI agents the ability to...  ..., building a groundbreaking platform that transforms how teams...  ...done.Forward is looking for a Site Reliability EngineerAbout the Role This...  ...platform. You will work closely with engineering, infrastructure,... 
    Platform
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  •  ...you want to think out of box with thriving on challenges in AI industry...  ...a System Software Engineer Lead, you willLead the team in...  ...Memory subsystem, coherency, AI/ML architecture, security, etc.Ability...  ...provide you the best possible platform to do that.Self-directed: We... 
    Platform

    Baidu

    Sunnyvale, CA
    4 days ago
  •  ...Staff Data Engineer, Full StackPrimary Skills: Data Engineering...  ...(Expert), Google Cloud Platform – BigQuery & Vertex AI (...  ...(Expert)Contract Type: W2 OnlyDuration: 9+...  ...platforms while driving AI/ML initiatives that deliver...  ...applications and AI agents.Partner with Product, GTM, Customer... 
    Platform
    Contract work
    Remote work

    Akraya

    Santa Clara, CA
    7 hours ago
  •  ...W2 Position Position: Data Engineer (PySpark/Python/ML Infrastructure) Location: CA, WA, NJ, NY Requirement: Hiring a Data Engineer with strong experience in PySpark, Python, SQL, and large-scale data pipelines. This role focuses on building... 
    Local area

    Octans Group LLC

    Sunnyvale, CA
    4 days ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE...  ...and improve service SLOs. We partner with Service Owners to drive reliability of...  ...design consulting, developing software platforms and frameworks, capacity management and... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $267k - $356k

     ...Tuesday.Lambda's Storage Engineering team is the backbone...  ...of Lambda's data platform services—from low-level...  ...industry, which means reliability and performance aren't...  ...automation and tooling.Partner with Storage Engineers,...  ...new and existing sites using tools such as Ansible... 
    Platform
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $101k - $161k

     ...networking to provide our clients with a competitive edge in an...  ...awards, such as Best Engineering Team, Best Company for Diversity...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... 
    Platform

    Arista Networks

    Santa Clara, CA
    5 days ago
  • $174k - $252k

     ..., developing software platforms and frameworks, capacity...  ...changes that improve reliability and velocity.Practice...  ...in Computer Science, Engineering, a related field, or equivalent...  ...5 years of experience with software development...  ...or Engineering.Site Reliability Engineering... 
    Platform

    Google

    Sunnyvale, CA
    2 days ago
  •  ...is currently Tuesday.Engineering at Lambda is responsible...  ...cloud networking platform and SDN infrastructureOperate...  ...with software, platform, and...  ...teams to improve service reliability and deployment workflowsDeploy...  ...years of experience in Site Reliability Engineering... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  •  ...accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our...  ...based in our Santa Clara, CA office, with an in-office schedule of two days per...  ...configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs... 
    Platform
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $90k - $180k

     ...technologies spans the spectrum of healthcare, with leading businesses and products in...  ...0 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our...  ...excellence of Merlin.net — a remote monitoring platform designed to help doctors,... 
    Platform
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $128.6k - $184.9k

     ...that powers our global cloud platform. As a team of six engineers distributed across the US, Canada...  ...deep infrastructure expertise with a strong focus on automation, reliability, and operational excellence....  ...Qualifications7+ years of experience in Site Reliability Engineering,... 
    Platform
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Santa Clara, CA
    1 day ago
  • $160k - $240k

     ...of times a day - quickly, reliably, and securely. Any time...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat...  ...successful Site Reliability Engineer do at Fiserv?You will...  ...and help operate financial platforms at scale. You will partner with cross-functional teams to... 
    Platform
    Full time

    Fiserv

    Sunnyvale, CA
    15 hours ago
  • $167.7k - $245.2k

     ...approximately 2 days per week on-site at Cisco offices in...  ...as intended, improving reliability and reducing risks....  ...-powered applications with enhanced observability...  ...Senior Site Reliability Engineer (SRE), you will build,...  ...Observability's deployment platform and production... 
    Platform
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Sunnyvale, CA
    6 days ago
  • $272k - $431.25k

     ...Principal System Software Engineer to drive next-...  ...innovations in automotive platform software, system architecture...  ....You will work closely with hardware, architecture,...  ...improve performance, reliability, determinism, and...  ...accelerated computing and AI/ML software platforms.... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges...  ...-on experience with observability platforms (e.g., Prometheus, Grafana).Strong... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...on GB300 Blackwell GPUs with NVLink interconnect,...  ...bandwidth architecture of this platform.We are looking for a...  ...systems software engineer who will own AI stack readiness...  ...DGX Station performs reliably as a shared workstation...  ...-on experience in AI/ML workload optimization,... 
    Platform
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...FlashDreams is our foundational platform for the next generation...  ...world models. We work with the top interactive world...  ...of research and systems engineering, where we turn new ideas into reliable, high-performance systems...  ...vision, or large-scale ML systems.Experience building... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $210.6k - $305.1k

     ...and operates our US GovCloud platform. This team is responsible for...  ...processes. Collaborate closely with cross-functional teams, including...  ...led a distributed team of 5+ engineers, can demonstrate strong...  ...Please see the Cisco careers site to discover more benefits and... 
    Platform
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    4 days ago
  • $224k - $356.5k

     ...medical devices. Our software platforms are central to this mission. We...  ...a Senior Systems Software Engineer to join our team as a technical...  ...maintain strong collaboration with automotive OEMs, robotics colleagues...  ...the Crowd:Experience with ML compiler frameworks (TVM, MLIR... 
    Platform
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    2 days ago
  • $100k - $105k

     ...Edge Computing AI Engineer – Remote Bright...  ...Full-time, Direct W2 Salary Range: $1...  ..., including mobile platforms, embedded systems,...  ...optimization, along with strong systems engineering...  ...skills to ship reliable AI capabilities outside...  ...of experience in ML engineering, with... 
    Platform
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Santa Clara, CA
    23 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer with ML platform - Only W2. Be the first to apply!