Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Observability & Telemetry Engineer [Remote]

Full-time

Bitdeer Technologies Group

San Jose, CA
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

We are seeking a Staff AI Observability & Telemetry Engineer to architect the "nervous system" of our AI-native NeoCloud platform. This role goes beyond standard monitoring; you are responsible for building the high-fidelity perception layer required to orchestrate massive-scale AI infrastructure. You will capture, store, and make sense of millions of hardware and software signals per second, enabling our SREs, automated remediation agents, and external customers to peer deep into the performance of their GPU workloads and the underlying network fabric. You will define the telemetry standards that drive our autonomous operations, ensuring we can detect, diagnose, and resolve hardware and software bottlenecks in real-time.

Key Responsibilities

  • Architect and scale a high-cardinality telemetry infrastructure using highly available time-series databases (e.g., VictoriaMetrics, Thanos, or Mimir) capable of handling massive ingestion rates.
  • Integrate complex hardware-level exporters (NVIDIA DCGM, network switch telemetry, IPMI/Redfish) directly into the Kubernetes observability stack to provide a unified view of the cluster.
  • Build eBPF-based diagnostic tools to trace network congestion, kernel-level I/O latency, and distributed training bottlenecks across the cluster.
  • Develop automated dashboards and alerting pipelines that trigger proactive cordoning of degraded hardware before it impacts customer training jobs.
  • Design the metric pipelines required for accurate, multi-tenant consumption billing based on real-time GPU and network utilization metrics.
  • Collaborate with the GPU Systems and Scheduling teams to create observability standards for "AI-native" workloads, ensuring deep insight into job efficiency and resource utilization.
  • Lead technical design reviews for observability architecture, mentoring team members on best practices for high-performance telemetry collection and analysis.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of software or site reliability engineering, with deep, hands-on expertise in the Prometheus/OpenTelemetry ecosystem.
  • Advanced proficiency in Go and extensive experience writing custom Kubernetes metric exporters and operators.
  • Hands-on experience with kernel-level tracing tools (eBPF, BCC) and deep performance tuning of Linux systems.
  • Strong familiarity with AI hardware metrics (GPU power states, SM utilization, memory bandwidth) and high-performance network telemetry.
  • Proven track record of operating, debugging, and scaling large-scale telemetry stacks in high-performance computing or cloud environments.
  • Strong technical leadership skills; ability to influence architectural decisions and align cross-functional teams around observability standards.
  • Excellent communication skills, with the ability to translate complex system requirements into manageable engineering milestones.
  • Experience working in high-velocity, high-growth engineering environments is strongly preferred.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 27 days ago
Similar jobs that could be interesting for youBased on the Staff AI Observability & Telemetry Engineer [Remote] in San Jose, CA vacancy
  • $150.98k - $218.62k

     ...Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions...  .... Learn more at and on LinkedIn and X.Staff System Architecture and Design...  ...DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Shift work
    Day shift

    Analog Devices

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're looking for a strong technical...  ...execution. ~ BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience). ~15+ years in... 
    Suggested
    Shift work

    NVIDIA Corporation

    Santa Clara, CA
    12 hours ago
  • $110k - $160k

     ...overall code quality. Own test reporting and articulately communicate technical challenges, solutions, and mitigation plans to both engineering teams and management.   Education / Experience: ~ Bachelors in Electrical Engineering, Digital Sciences, Computer... 
    Suggested

    SK hynix memory solutions America Inc.

    San Jose, CA
    19 days ago
  • $2,500 per month

     ...investors and staffed by leading engineers, Etched is redefining the...  ...chip design, simulation, and AI model deployment in a cloud and...  ...design workflows. Build observability solutions using Grafana, Prometheus...  ...expect all of our technical staff to contribute to both and... 
    Suggested
    Work at office
    Relocation package

    Etched

    San Jose, CA
    10 days ago
  • $110k - $175k

     ...computing.   About the Role : As a Firmware Validation Engineer at SK hynix memory solutions, you will drive enterprise SSD (eSSD...  ...languages like JavaScript or python. Good understanding of AI development tools. Fast learner with good communication skills... 
    Suggested

    SK hynix memory solutions America Inc.

    San Jose, CA
    2 days ago
  • $136k - $218.5k

     ...Group is seeking a versatile engineer to join the HW ArchDev Methodology...  ..., automation frameworks, and telemetry systems that determine whether next-generation GPUs and AI accelerators ship on time and...  ...others still use: pipelines, observability systems, debug tools You... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $154k - $200k

     ...seeking a self-driven, hands-on Process Engineer to help scope, define, and design balance...  ...thermal energy storage systems. Apply AI tools thoughtfully to improve engineering...  ...features flexible and inclusive holiday observance, as well as paid volunteer time off. #LI... 
    Flexible hours

    Antora Energy

    San Jose, CA
    9 days ago
  • $152k - $287.5k

     ...now looking for a Senior System Software Engineer to join our Robotics Team, with a strong...  ...system software, optimized robotics and AI algorithms, and powerful edge computing...  ...validation, testing, benchmarking, and observability infrastructure spanning simulation, hardware... 
    Full time

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...Description Job Description About Boson AI: At Boson AI, we are not just building...  ...collaborative team of researchers and engineers who thrive on pushing the boundaries of...  ...retrieval pipelines. Instrument end-to-end observability: define SLIs/SLOs, build structured... 

    Boson AI

    Santa Clara, CA
    5 days ago
  • $147.9k - $220k

     ...seeking a Senior Network Systems Engineer to join our Global Network Services...  ...Maintain operational health through observability — use network monitoring and telemetry tooling to detect, diagnose, and...  ...Model Context Protocol (MCP) and AI-agent integration — connecting AI... 
    Permanent employment
    Work at office
    Local area
    Worldwide

    NetApp

    San Jose, CA
    3 days ago
  • $186.2k - $273.07k

     ...quality standards for technologies like AI, data centers, automotive and electronic...  ...experienced Senior eBeam Systems Design Engineer to lead the design, modeling, optimization...  ...simulation tools and real‑world system observations to improve performance, reliability, and... 
    Minimum wage
    Work experience placement

    KLA

    Milpitas, CA
    2 days ago
  • $266.05k - $396k

     ...help organizations unlock the full potential of their data, from AI to multicloud. Ready to innovate and contribute to our path to $...  ...excel together across every function. Job Summary Distinguished Engineer - AI Infrastructure We are seeking a Distinguished Engineer with... 
    Work at office
    Local area

    NetApp

    San Jose, CA
    3 days ago
  • $124k - $155k

     ...a Great Place to Work, we offer a deep observability pipeline that efficiently delivers network...  ...organizations. As a Senior SW QA Engineer on the GigaSMART team, you will lead quality...  ...We may use automated tools, including AI-based systems, to help screen and evaluate... 
    Local area
    Worldwide

    Gigamon

    Santa Clara, CA
    28 days ago
  • $160.1k - $285.1k

     ...As a Fortune 500 company and a leading AI platform for managing people, money, and...  ...About the Team The Data Platform and Observability team is based in Pleasanton, CA; Boston,...  ...you want to work alongside world-class engineers building the platforms that make that possible... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Santa Clara, CA
    4 days ago
  • $165k - $235k

     ...RoboForce RoboForce is an AI robotics company developing Physical...  .... The company’s robots are engineered for demanding industrial...  ...velocity, testing efficiency, observability, and release quality across...  ...tools for testing, simulation, telemetry, debugging, or log analysis.... 
    Full time
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    4 days ago
  • $190k - $219k

     ...Muon seeks a Senior Mission Lead Systems Engineer to join our Mission Engineering team. At...  ...system modeling, develop test plans with AI&T teams, and identify and reduce risk across...  ..., analysis, and trade spaces for Earth observation missions. Experience with space system... 
    Permanent employment
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Muon Space

    San Jose, CA
    12 hours ago
  • $183k - $247.6k

     ...-accelerated servers powering AI/ML training and inference at cloud...  ...a Cloud Hardware Development Engineer to define server architectures...  ...modes* Monitor operational telemetry to identify systemic issues and...  ...employees, supervisors, and staff; adhere to standards of excellence... 
    Local area
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    3 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    3 days ago
  • $44k - $185k

     ...Team The Cisco Product Organization is the engine behind the platforms, technologies, and...  ...generation of connectivity. Spanning AI Software & Platforms, Core Product Engineering...  ..., Security, Networking, Silicon, and Observability, we build the critical infrastructure that... 
    Full time
    Temporary work
    Work experience placement
    Summer work
    Internship
    Local area
    Flexible hours

    Cisco

    Milpitas, CA
    1 day ago
  •  ...years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or...  ...environments. Knowledge of monitoring, observability, and reliability engineering practices...  ...platforms. Experience leveraging AI-assisted development tools to improve software... 
    Full time
    Worldwide

    SFE

    San Jose, CA
    2 days ago
  • $110k - $175k

     ...wide range of distinguished customers globally. Description We are looking for a driven, team-oriented Senior Signal Integrity Engineer to develop next generation Solid State Drive (SSD) products. This position requires experience in hardware design and signal and... 
    Work experience placement
    Flexible hours

    SK hynix memory solutions America Inc.

    Santa Clara, CA
    19 days ago
  • $116k - $159.5k

     ...Applied Materials is the global leader in materials science and engineering solutions that are at the foundation of virtually every new...  ...equipment that we create and service is essential to advancing AI and accelerating the commercialization of next-generation semiconductor... 
    Full time
    Relocation

    APPLIED MATERIALS

    Santa Clara, CA
    6 hours ago
  • $195k - $225k

     ...talented, passionate, and committed engineers, technologists, and business...  ...to join us.Job Summary:As a Staff Data Center Solutions...  ...particular focus integrating AI Cloud Systems monitoring and application...  ...well-instrumented control and telemetry systems.Essential Duties and... 
    Worldwide

    Supermicro

    San Jose, CA
    12 hours ago
  • $122.5k - $175k

     ...believe the future of work is Human + AI and are building an AI-native...  ...at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role...  ...applicationsMaintain platform security and observability through nftables and comprehensive... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    4 days ago
  • $141.3k - $360.7k

     ...state-of-the-art technology and engineering, keeping the customer at the...  ...do that work with agentic AI as your default mode of...  ...in the right code, schemas, telemetry, and business logic, and keep...  ...Establish evaluation, observability, and monitoring for the signals... 
    Full time
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $200k - $322k

    NVIDIA is looking for a skilled and motivated Senior Infrastructure Engineer to join our dynamic team. You will contribute to innovative...  ...stand out from the crowd:Proficiency in brand new technologies like AI, machine learning, and data-driven solutions.Proven ability to... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $50 - $70 per hour

     ...Infrastructure Engineer (Linux, Datacenter & Automation) Location: Santa Clara, CA...  ...infrastructure environments. Exposure to AI, HPC, GPU, cloud, or accelerated...  ...environments. Experience with monitoring and observability platforms such as Prometheus, Grafana,... 
    Contract work
    Temporary work

    TEKsystems

    Santa Clara, CA
    4 days ago
  • $203k - $270k

     ...allies. Founded by two former U.S. Navy electrical engineers with deep experience in robotics and software, ACS brings together AI, computer vision, precision motion, and...  .... About The Role: We are looking for a Staff Level Computer Vision and Machine Learning Engineer... 
    Local area

    Allen Control Systems

    Mountain View, CA
    4 days ago
  • $176.1k - $308.2k

     ...de l'entreprise It all started when engineer Fred Luddy wrote code that automated a tedious...  ...meaningful work. Today, ServiceNow is the AI control tower for business reinvention....  ...About the role  We are looking for a Staff FinOps AI Governance Lead to drive financial... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    5 days ago
  • Sandisk seeks an experienced system-level test engineer to lead development and qualification of test solutions for advanced NAND technologies. You will drive cross-functional efforts across Product and Test Engineering, align with customer expectations, and ensure high... 

    Sandisk

    Milpitas, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Observability & Telemetry Engineer [Remote]. Be the first to apply!