Staff AI Observability & Telemetry Engineer
Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
We are seeking a Staff AI Observability & Telemetry Engineer to architect the "nervous system" of our AI-native NeoCloud platform. This role goes beyond standard monitoring; you are responsible for building the high-fidelity perception layer required to orchestrate massive-scale AI infrastructure. You will capture, store, and make sense of millions of hardware and software signals per second, enabling our SREs, automated remediation agents, and external customers to peer deep into the performance of their GPU workloads and the underlying network fabric. You will define the telemetry standards that drive our autonomous operations, ensuring we can detect, diagnose, and resolve hardware and software bottlenecks in real-time.
Key Responsibilities
- Architect and scale a high-cardinality telemetry infrastructure using highly available time-series databases (e.g., VictoriaMetrics, Thanos, or Mimir) capable of handling massive ingestion rates.
- Integrate complex hardware-level exporters (NVIDIA DCGM, network switch telemetry, IPMI/Redfish) directly into the Kubernetes observability stack to provide a unified view of the cluster.
- Build eBPF-based diagnostic tools to trace network congestion, kernel-level I/O latency, and distributed training bottlenecks across the cluster.
- Develop automated dashboards and alerting pipelines that trigger proactive cordoning of degraded hardware before it impacts customer training jobs.
- Design the metric pipelines required for accurate, multi-tenant consumption billing based on real-time GPU and network utilization metrics.
- Collaborate with the GPU Systems and Scheduling teams to create observability standards for "AI-native" workloads, ensuring deep insight into job efficiency and resource utilization.
- Lead technical design reviews for observability architecture, mentoring team members on best practices for high-performance telemetry collection and analysis.
Qualifications
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of software or site reliability engineering, with deep, hands-on expertise in the Prometheus/OpenTelemetry ecosystem.
- Advanced proficiency in Go and extensive experience writing custom Kubernetes metric exporters and operators.
- Hands-on experience with kernel-level tracing tools (eBPF, BCC) and deep performance tuning of Linux systems.
- Strong familiarity with AI hardware metrics (GPU power states, SM utilization, memory bandwidth) and high-performance network telemetry.
- Proven track record of operating, debugging, and scaling large-scale telemetry stacks in high-performance computing or cloud environments.
- Strong technical leadership skills; ability to influence architectural decisions and align cross-functional teams around observability standards.
- Excellent communication skills, with the ability to translate complex system requirements into manageable engineering milestones.
- Experience working in high-velocity, high-growth engineering environments is strongly preferred.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$137.04k - $188.43k
We’re looking for a Staff Integration & Application Developer to... ...specifically leverage Workato AI Agents to accelerate the... ...Rules" for governance.Automated Observability: Implement AI-driven job monitoring... ...background in integration engineering, specifically within the...SuggestedFull timeWork at officeLocal areaImmediate start$183k - $247.6k
The Annapurna AI Manufacturing, Quality and Reliability (MQR) Team is part of AWS Annapurna Labs focused... ...Cloud Services provider. As a Senior Reliability Engineer you will engage with an experienced cross-disciplinary staff to conceive and design infrastructure...SuggestedWork experience placementLocal areaFlexible hours$173.9k - $235.2k
...you want to shape the future of Generative AI at AWS? Join the team building the... ...industry.You’ll join a diverse AWS Hardware Engineering team of software, hardware, and network engineers... ...with other employees, supervisors, and staff; adhere to standards of excellence...SuggestedInternshipLocal areaFlexible hours- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded... .... Together, we advance your career. THE ROLE:Analysis (FA) Engineer, you will play a critical role in diagnosing, isolating, and resolving...Suggested
$173.9k - $235.2k
...experienced Senior Systems Development Engineer to lead the development of... ...our edge and accelerated (AI/ML) compute fleet healthy —... ...failure detection systems using telemetry, sensor data, error trending,... ...other employees, supervisors, and staff; adhere to standards of...SuggestedInternshipLocal areaWorldwideFlexible hours$183k - $247.6k
Do you want to shape the future of AI? Join the team building the foundation of the world... ...team of software, hardware, and network engineers, supply chain specialists, security... ...cooperatively with other employees, supervisors, and staff; adhere to standards of excellence...Local areaFlexible hours- ...Senior/Staff Computer Vision & Machine Learning Engineer for Autonomous Anti-Drone Systems Company Overview: Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software...Full timeLocal area
$148.7k - $201.2k
...build the backbone of Generative AI at AWS? Do you want to build... ...seeking a Systems Development Engineer to develop automation software... ...detection systems using telemetry, sensor data, error trending,... ...other employees, supervisors, and staff; adhere to standards of excellence...InternshipLocal areaWorldwideFlexible hours- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded... ...your career. The Role AMD’s Partner Support Field Applications Engineering Organization is responsible for supporting our OEM SA with end...
- ...characterization and analysis, we influence future SoC, CPU and IP designs. We are seeking highly skilled and motivated simulation/emulation engineers to join our diverse team at Arm!Responsibilities:As a SoC verification/Emulation Engineer, you will:Develop SoC verification...Work at officeLocal areaRemote work
- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems.... ...career. THE ROLE: We are looking for a dynamic, energetic Lead / Staff Systems Engineer to join our growing team. As a key contributor to the success...
- ...Fluidstack Production Engineering Team Examples of key exciting problems the... ...GPUs legible in real time: build the observability platform that turns raw telemetry into signal, from site-level... ...and move on. You're fluent with AI tooling. LLM APIs, MCP servers, and...Local area
- Apple Inc. in Austin, Texas, seeks a Senior Software Engineer specialized in observability platforms. In this role, you will design and develop systems that enhance visibility into service performance, ensuring reliability and security. With over 7 years of experience in...
$88.97k - $287.91k
...Forward Deployed Engineer (FDE) Focusing on Agentforce Operations As a Forward Deployed... ...and deploys enterprise-grade agentic AI solutions directly inside customer... ...deployable code in days, not months. Telemetry & Observability: Model data, ship ETL/ELT pipelines,...- ...services provider in Austin, Texas, seeks a Senior Performance Engineer to drive in-sprint performance engineering for complex,... ...NFRs into user stories, test authoring, performance tuning, and observability. This role offers an exciting opportunity to engage in automated...
$203k - $270k
...Job Description Job Description Staff Computer Vision & Machine Learning Engineer for Autonomous Anti-Drone Systems Company Overview: Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track...Local area- ...As a senior Machine Learning Systems Engineer on the Search Platform team, you will own... ...retrieval workflows. Partner with Rovo and AI platform teams to evolve search... ...Operational Excellence & Cost DisciplineDrive observability, monitoring, and incident response for search...Work at officeLocal area
- AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting Engagement)Employment... ...RoleWe are seeking an experienced AI/ML Engineer to design, deploy, and operate... ...requiring high availability, scalability, and observability.Compensation and Benefits Slalom prides...Full timeContract workLocal areaRemote workFlexible hours
- ...to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever... ...or equivalent middleware on real hardware. You engineer for remote failure: recovery, fallback, and observability built in. You've integrated robot systems with...Full timeRemote work
- ...to spot a "Normalized" problem and the AI-native curiosity to create a solution using... ...many providers, with reliability, observability, security, cost control, and routing built... ...change. We are looking for a Senior Systems Engineer to help build that layer. This is a...Local area
$98k - $130.8k
...looking for a highly motivated L2 Support Engineer with strong automation expertise to join... ...-critical platforms. Experience with AI-powered tools and intelligent automation... ...Hands-on experience with monitoring and observability tools.Strong troubleshooting skills in high...Permanent employmentImmediate startRemote work- ...world's top brands, offering comprehensive engineering, supply chain, and manufacturing... ...programming and scripting, open-source tooling, AI-driven automation, low-code platforms,... ...(Proxmox, VMware, or similar); observability and monitoring (Grafana, Prometheus); containerization...Full timeWork at officeLocal areaWorldwide
- ...We are seeking highly skilled and motivated performance analysis engineers to join our diverse team at Arm!Responsibilities:As a... ...accelerate performance analysis and root-cause identification.Leverage AI-based tools and automation to improve engineering productivity and...Work at officeLocal areaRemote work
$250k - $300k
...things they had to do. Powerful AI will be the biggest lever... ...the frontier moving. Bad telemetry is a trust failure at... ...setting the data contracts and engineering standards other teams build... ...instrumented production services with observability tooling (Prometheus, Grafana...Full time$140k - $215k
...with the world’s most advanced AI-native platform. We work on... ...installed on client machines that observes system activity. When the... ...capability and remote telemetry to the Falcon cloud. The cloud... ...on Linux platforms. Our tools engineers own development of our test pipelines...Full timeWork experience placementWork at officeLocal areaRemote work2 days per week3 days per week$120k - $180k
...with the world’s most advanced AI-native platform. We work on... ...installed on client machines that observes system activity and... ...prevention capability and remote telemetry to the Falcon Host cloud. The... ...their environments.This is an Engineer 3 - macOS Engineer role in the...Full timeWork experience placementWork at officeLocal areaRemote work2 days per week3 days per week- ...digital core and unleashing the power of AI to create value at speed. With 779,000 professionals... ...technology meets real‑world impact. Our engineers, architects, and security specialists... ...such as for a disability or religious observance, please call us toll free at 1 (877) 889-...Full timeWork experience placementLive inWork at officeLocal areaWorldwide
$202k - $310k
...DescriptionThe RoleAs a Lead System Integration Engineer - Enterprise Transformation, you will... ...Workspace, SaaS platforms, enterprise AI tools, and GM’s core business systems through... ..., and maintain robust, resilient, and observable integrations using APIs, webhooks,...Full timeH1bLocal areaWork from homeRelocation package- Oracle is seeking a Senior Staff Engineer to shape OCI infrastructure, lead multi-team initiatives... ...for reliability, security, and observability. You will mentor engineers, drive architectural... ...languages, and experience with AI-assisted development. #J-18808-Ljbffr...
$233.99k - $330.34k
...Description: We are seeking an innovative and experienced engineering leader to join our developer software team. In... ...Plus, you have the opportunity to transform how observability, debugging and profiling is done in the modern AI era, enabling engineers internally and...Full timeTemporary workWork experience placementLocal areaImmediate startShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff AI Observability & Telemetry Engineer. Be the first to apply!
- staff engineer Austin, TX
- assistant engineer Austin, TX
- research assistant engineering Austin, TX
- staff design engineer Austin, TX
- engineering aide Austin, TX
- senior staff engineer Austin, TX
- senior staff systems engineer Austin, TX
- software engineer staff Austin, TX
- technology administrator Austin, TX
- assistant mechanical engineer Austin, TX



