Tech Lead - Network Observability
$180k - $260kClockwork Inc
About Clockwork Systems Clockwork.io - Software Driven Fabrics to increase GPU cluster utilization Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of distributed computing. As AI workloads grow increasingly complex, traditional infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion.
To learn more, visit About the Role We are seeking an experienced Tech Lead to lead the architecture, development, and scaling of a high-performance network monitoring and observability platform. This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks. The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus).
What You Will Do
Compensation for this position will vary based on the skills and experience you bring, as well as internal equity considerations. For candidates hired at the posted level, the expected base salary range is $180,000 - $260,000. The offered compensation package may also include stock options or other equity awards, subject to Clockwork's equity program and applicable approvals In addition to cash compensation, this role is eligible to participate in the company's equity program, which may include stock options granted in accordance with the company's equity plan and subject to approval and applicable vesting schedules. Clockwork Systems is an equal opportunity employer. We are committed to building world-class teams by welcoming bright, passionate individuals from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, age, sex, sexual orientation, gender identity or expression, national origin, disability, or protected veteran status. We believe diversity drives innovation, and we grow stronger together.
To learn more, visit About the Role We are seeking an experienced Tech Lead to lead the architecture, development, and scaling of a high-performance network monitoring and observability platform. This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks. The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus).
What You Will Do
- Lead architecture, design, and development of scalable network monitoring platforms for high-performance RDMA, RoCE, InfiniBand, and TCP/IP infrastructure.
- Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows.
- Troubleshoot complex production issues across application, OS, server, RDMA, and network layers while optimizing low-latency collection, aggregation, and alerting.
- Establish engineering standards, drive automation, define technical roadmaps with cross-functional teams, and mentor engineers on distributed systems and high-performance networking best practices.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages.
- Proven experience leading engineering teams, major technical initiatives, or complex infrastructure projects.
- Experience building distributed systems, backend services, telemetry pipelines, or observability platforms.
- Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics.
- Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, memory registration, and related RDMA concepts.
- Strong knowledge of Linux networking, TCP/IP, DNS, routing, MTU, congestion control, packet loss, latency, and performance tuning.
- Experience with traceroute-style diagnostics, path discovery, network reachability checks, synthetic probes, or active network measurements.
- Experience with monitoring and visualization platforms such as Prometheus, Grafana, Datadog, Splunk, OpenTelemetry, or similar tools.
- Strong debugging skills across software, operating system, server, and network layers.
- Experience operating production systems in Linux-based environments.
- Strong architectural judgment and ability to design systems for reliability, scalability, and operational simplicity.
- Experience supporting AI/ML, HPC, storage, or GPU cluster infrastructure workloads.
- Experience with large-scale RoCE or InfiniBand deployments.
- Experience with NCCL, distributed training infrastructure, or AI cluster diagnostics.
- Experience with eBPF, XDP, DPDK, perf, tcpdump, Wireshark, ethtool, iproute2, rdma-core, or Linux kernel networking tools.
- Experience with cloud infrastructure on AWS, GCP, or Azure.
- Experience with Kubernetes, service discovery, configuration management, and infrastructure automation.
- Knowledge of security, compliance, and infrastructure best practices.
- Experience designing time-series data systems, alerting pipelines, or high-cardinality telemetry platforms.
- Challenging projects.
- A friendly and inclusive workplace culture.
- Competitive compensation.
- A great benefits package.
- Catered lunch.
Compensation for this position will vary based on the skills and experience you bring, as well as internal equity considerations. For candidates hired at the posted level, the expected base salary range is $180,000 - $260,000. The offered compensation package may also include stock options or other equity awards, subject to Clockwork's equity program and applicable approvals In addition to cash compensation, this role is eligible to participate in the company's equity program, which may include stock options granted in accordance with the company's equity plan and subject to approval and applicable vesting schedules. Clockwork Systems is an equal opportunity employer. We are committed to building world-class teams by welcoming bright, passionate individuals from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, age, sex, sexual orientation, gender identity or expression, national origin, disability, or protected veteran status. We believe diversity drives innovation, and we grow stronger together.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Tech Lead - Network Observability in Palo Alto, CA vacancy
- ...of millions of virtual machines, generating terabytes of logs and processing exabytes of data per day. At our scale, we observe cloud hardware, network, and operating system faults, and our software must gracefully shield our customers from any of the above. As a...NetworkFull time
- ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments...SuggestedFull time
- ...infrastructure used for hardware-in-the-loop (HIL) testing. Monitor the overall health of test devices and infrastructure; perform basic network/power troubleshooting, maintain device logs, and escalate complex hardware issues. Configure and manage embedded devices via...NetworkLocal areaRemote work
$180k - $260k
...AI fabrics by delivering cross-stack observability to catch and quickly resolve problems,... ...looking for a passionate and experienced Tech Lead - Frontend / Full Stack to join our... ...and turning complex infrastructure and network data into clear, intuitive visual experiences...Network- ...1.5 million users across 10,000+ teams. Leading brands—including Alibaba, HubSpot, Binance... ...operating-systems–level software, compilers, and network distribution software for massive social... ...in production. ~ Prior experience in a tech lead (TL) capacity, such as driving...NetworkFull time
$240k - $400k
...on, customer facing delivery. You will lead builds across Node.js services, AI and agent... ...secure, scalable systems across networking, autoscaling, multi tenant patterns. Proficiency... ...Code using Terraform or CDK and strong observability with metrics, tracing, logs, and SLOs....NetworkFull timeVisa sponsorship$140k - $210k
...to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems,... ...who will design the next generation of network acceleration software. Stay a nanosecond... ...Use your knowledge and experience to lead/contribute directly to the design and build...Network- ...security, and reliability. Build and maintain monitoring and observability solutions that support application performance,... ...Strong understanding of IaaS and PaaS offerings, IAM, and networking within cloud environments. Infrastructure as Code (IaC): Proficiency...Network
- ...to inference/robot endpoints. Improve our infrastructure observability. Harden security & multi-tenancy Qualifications ~5+ years... ..., gRPC, etc.) Experience with ML Ops Understanding of networking fundamentals (TCP/IP, DNS, firewalls) and security best...Network
$115k - $230k
...traditional IT model to a modern tech organization with engineering... ...standards. ~ Lead the design and architecture of... ...~ Experience working with networking, caches, key/value stores, load... ...platforms (e.g., Kubernetes), and observability. ~ Experience with infrastructure...NetworkHourly payWork experience placementLocal area$152k - $248k
...LinkedIn is the world's largest professional network, built to create economic opportunity... ...SLOs/SLIs, instrument services for observability (metrics, logging, tracing), reduce toil... ...decisions. Participate in on-call rotation, lead incident response, and drive blameless...NetworkFor contractorsWork experience placementWork at officeFlexible hours$117.2k - $176.7k
...Here, ambition meets action. Tech meets trust. And innovation isn... ...your career at the company leading workforce transformation in the... ..., policy enforcement, and observability improvements, with guidance from... ...workloads or clusters (networking, scaling, upgrades). Good programming...NetworkFull time$228k - $242k
...designed to be easy to integrate into existing delivery and logistics networks, offering a scalable drone delivery solution for a broad range... ...a highly motivated and experienced Reverse Logistics Technical Lead to join our Global Central Operations team in Palo Alto, CA . In...NetworkFull timeLocal area$73.8k - $218.8k
...— our federated Model Context Protocol network across Oracle product lines, including retrieval... ...01,300 About Accenture Accenture is a leading global professional services company... ...such as for a disability or religious observance, please call us toll free at 1 (877) 889...NetworkFull timeWork experience placementLive inWork at officeLocal area$183k - $253k
...exchange between connected devices, reducing latency and improving network efficiency through peer-to-peer connectivity Collaborate... ...and third-party aggregators Experience with monitoring and observability tools for tracking system performance, analyzing API usage...NetworkFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift$297.32k
...including service boundaries, data flows, observability patterns, and modernization of legacy... ...customer facing product experiences Lead technical delivery across the full lifecycle... ...on AWS, including compute environments, networking, container orchestration, storage, and...NetworkFull timeWorldwide2 days per week3 days per week$152k - $228k
...? How does a new data format impact onboard logging rate or network contention as more data might be flowing from through the system... ...and release-blocking quality gates. Platform Reliability & Observability: Build monitoring, alerting, and self-healing automation for...NetworkFull timeTemporary work- ...Senior Lead Site Reliability Engineer Elevate your engineering prowess to unprecedented... ...with others to create and implement observability and reliability designs for complex... ...engineering community, and continues to expand network and leads evaluation sessions with...Network
- ...Lead Site Reliability Engineer Assume a critical role in defining the future of a... ...Implements infrastructure, configuration, and network as code for the applications and... ...within applications and platforms; strong observability background including white/black-box monitoring...Network
- ...echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan... ...Collaborates with others to create and implement observability and reliability designs for complex... ...community, and continues to expand network and leads evaluation sessions with vendors...Network
$143.1k - $226k
...leader in cloud infrastructure, data center networking, and security, is seeking a Core... ...customers to shape the future of our industry‑leading platform. Key Responsibilities... ...and Kubernetes networking Knowledge of observability tools such as Prometheus and Grafana Contributions...NetworkLocal area$265k - $331.3k
...enterprise outcomes. Responsibilities Lead the technical design and hands‑on... ...reviews; improve code quality, reliability, observability, and cost/performance of AI workloads.... ...professional, employment, social media/website, network/device, recruiting system usage/...NetworkFull timeContract workLocal area$170k - $277k
...Palo Alto Networks, Inc. is seeking a Principal Software Engineer to drive the technical leadership of innovative cloud security solutions. The ideal candidate will have over 15 years of experience in software engineering and a strong command of programming languages...NetworkFull time- ...Overview Software Engineer (Networking & Telemetry Systems) at Microsoft. Join to apply... .... We are seeking a Software Engineer to lead the design and development of scalable networking... ...infrastructure, with a focus on robust observability and debugging capabilities....NetworkFull time
$210.3k - $273.4k
...server compatibility in mind, and drive observability across system health and business metrics... ...issues like performance regressions and network latency through strong observability and... ...at Unity Unity [NYSE: U] is the world’s leading game engine, powering play for more than...NetworkTemporary workWork at officeWorldwideRelocation package$209.6k - $314.4k
...field, where "reinstall the machine" is not an option. Own the Network Between Cloud and Printer: Design the protocols, connections,... ...mindset: you think about failure modes, blast radius, observability, and upgrade paths from the start, and you have carried a pager...NetworkPermanent employmentFull timeFor contractorsWork at officeRemote work$2,000 per month
...based solutions for search, security, and observability help organizations deliver on the... ...pipelines reliable and secure. Security & Networking: Apply strict security and network... ...Platform Landscape: Familiarity with leading agentic AI and workflow automation platforms...NetworkLocal areaFlexible hours- ...based execution pipelines, GPU/CPU workload optimization, system observability, and performance engineering for production embedded systems.... ...including scheduling, memory management, concurrency, IPC, networking, logging, and storage. Strong understanding of distributed...NetworkFull timeLocal areaWorldwideFlexible hoursShift work
$125k - $175k
...RESPONSIBILITIES: Design, build, and scale a new Content Delivery Network (CDN) for Starlink Build robust, responsive, and scalable... ...systems which shape traffic across multiple CDNs. Participate in and lead architecture, design, and code reviews. Develop prototypes and...NetworkPermanent employmentFull timeTemporary workInternshipWork at officeWorldwideMonday to FridayWeekend work- ...Manager/Tech Lead, Network Engineering CTH On-site — Sunnyvale, CA About this role: We're looking for a hands-on Manager... ...expected to mature it into a globally consistent, secure, and observable network as we open new offices and grow headcount. What...NetworkWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tech Lead - Network Observability. Be the first to apply!
Related searches
- network intern Palo Alto, CA
- palo alto networks Palo Alto, CA
- networking Palo Alto, CA
- IT network Palo Alto, CA
- senior cloud network engineer Palo Alto, CA
- family healthcare network Palo Alto, CA
- cloud network engineer Palo Alto, CA
- data network cabling Palo Alto, CA
- computer network Palo Alto, CA
- network operations center technician Palo Alto, CA













