Get new jobs by email
$166k - $220k
...the military in months, not years.ABOUT THE TEAMOur DeviceOS team develops the operating system that powers Anduril’s robots. We build... ...Yocto, buildroot, or similar systemsFamiliarity with packaging CUDA libraries and applicationsFamiliarity with functional programming...SuggestedFull timeWork experience placementImmediate start$220k - $292k
...years.ABOUT THE TEAMThe Tactical Recon & Strike team at Anduril develops aerial small drones (Group 1-3) and all equipment to test, deploy... ...Yocto, buildroot, or similar systemsFamiliarity with packaging CUDA libraries and applicationsFamiliarity with functional programming...SuggestedFull timeWork experience placementImmediate start$272k - $431.25k
...security across a large engineering org.Deep familiarity with the operational and deployment aspects of the NVIDIA AI/ML software stack (CUDA, cuDNN, containerization).Patent contributions or a strong publication record in areas related to distributed systems, cloud...SuggestedFull time$184k - $287.5k
...from concept through deployment.What you’ll be doing:Design and develop software solutions for data center servers including Linux kernel... ...Ways to stand out from the crowd:Experience with GPU computing (CUDA), deep learning workloadsExpertise in Out of Band and In-band...SuggestedFull time$168k - $270.25k
...root causing sophisticated customer issuesWork with R&D teams to develop bug fixes, workarounds, and solutions for critical customers... ...HPC performance test tools and NVIDIA AI stacks (NCCL, MPI, DOCA, CUDA)Widely considered to be one of the technology world’s most desirable...SuggestedFull timeWeekend work- ...services and operating them on Kubernetes in production.Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.Demonstrated ability to harden complex systems for...Suggested
$224k - $356.5k
...validated recipes, and model bring-up infrastructure — that lets developers run groundbreaking LLMs out of the box. The open-source... ...Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — you understand how thread blocks,...SuggestedFull timeLocal area$141k - $225.6k
...Experience with hardware-accelerated video pipelines a plus.Hands-on experience with inference at the edge using NPUs or GPUs — TensorRT, CUDA, ONNX Runtime, or similar a plus.Benefits that Benefit YouCompetitive salary and 401k with employer matchDiscretionary paid time...SuggestedWork experience placementWork at officeRemote workWorldwide$184k - $287.5k
We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software... ..., Volcano, or similar projects.Experience with GPU computing (CUDA), deep learning workloadsNVIDIA is widely considered to be one...SuggestedFull timeRemote work$168k - $239k
...instrumentation for performance monitoring (CPU, GPU, latency, memory) and develop offline benchmarking frameworks, tools, and scripts to evaluate... ...or related field and 3+ years of experience.Strong knowledge of CUDA as applied to recent GPU microarchitectures (e.g., Ampere,...SuggestedFull timeTemporary workRelocation package$224k - $356.5k
We are seeking a Senior Developer Relations Manager to join our team. You'll build establishing relationships across developer communities... ...computing, AI, or GPU acceleration platforms, including CUDA, TensorRT LLM, Dynamo, Triton Inference Server, NeMo, Isaac Sim...SuggestedFull timeShift work- ...the firm’s portfolios. You will lead virtual and direct teams of developers, teaching them best practices in high-performance computing (... ...organizations and regulated industries is a plusHands-on experience with CUDA for GPU programming and performance optimization preferred....Suggested
$184k - $287.5k
...deep expertise to manage complex customer engagements and help develop our product and architecture direction. This role offers an outstanding... .../CSPs.Hands-on expertise with RDMA verbs, DPDK, DOCA, NCCL, CUDA-aware networking, congestion control, and performance tuning at...SuggestedFull timeShift work$135.2k - $306.4k
...GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, RoCE, InfiniBand, RDMA, SmartNIC/DPU offload, or high-... ..., testing, debugging, documentation, operational analysis, and developer productivity while maintaining strong ownership, security judgment...SuggestedTemporary workRemote workFlexible hours$152k - $241.5k
...any conditions? If so, join us!What You'll Be Doing:Design and develop algorithms for map-based driving productsArchitecture design for... ...Buffers and Protocol BuffersExperience with GPGPU programming (CUDA)We believe, realizing self-driving vehicles will be a defining contribution...SuggestedFull time$143.7k - $194.4k
...platform launches.Your impact will extend from low-level systems (CUDA, EFA, firmware) through ML frameworks to serving layers,... ...that balance customer requirements with operational excellence- Develop comprehensive regression test coverage across all major component...InternshipFlexible hoursNight shift- ...across the stack, in ways that raises the bar for the entire AMD AI developer experience.Develop reproducible benchmarks and performance... ..., and framework-level work.Experience with GPU programming via CUDA and/or HIP, including writing and optimizing compute kernels beyond...
$92k - $135k
...demo. Coursework/research with PyTorch or TensorFlow; simple CUDA projects a plus. Familiarity with Grafana/Prometheus/OpenTelemetry... ...that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$110k - $170k
CompanyFounded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the digital economy. Taxbit's AI-enabled platform streamlines compliance related to digital assets, payments, and other financial transactions. Its ...Work at officeWork from home$196k - $294k
...and deployment mission execution software. Your decisions will significantly impact the company today and in the future. We seek a developer who quickly understands distributed systems and creates features that drive bottom-line value for the business. WHAT YOU'LL...Full timeWork experience placementLocal areaRelocation package- The Role Join us in revolutionizing the lending landscape. SoFi is seeking enthusiastic Principal Software Engineers who are ready to lead the technical and strategic evolution of our financial services platform in support of our goals that put our members in control...Full timeTemporary workWork experience placement
$184k - $287.5k
...and differentiable programming on CPUs and NVIDIA GPUs. Our team develops Warp for computational physics, AI, and optimization workflows.... ...adopting Warp in computational applications, from native C++ and CUDA changes through production integration, validation, and long-...Full time$168.1k - $227.4k
...breakthrough performance optimizations for AWS Trainium and Inferentia• Develop ML tools to enhance LLM accuracy and efficiency• Transform... ...benchmarking of AI accelerators- Experience in developing CUDA kernels, HPC and inference optimization, tensors operationsAmazon...Work experience placementWork from homeFlexible hours$143.7k - $194.4k
...leverage both your engineering and machine learning background to help develop generative AI for shopping. On a day-to-day basis, you will:-... ...management, and parallel computing principles- Experience with CUDA/C++/Kernel developmentAmazon is an equal opportunity employer...InternshipFlexible hours$184.5k
At Expedia Group, we help travelers explore the world, one journey at a time. As a global travel company powered by passionate people, trusted partnerships, and leading technology, we connect travelers, partners, and advertisers through our consumer brands, B2B network,...Full time$226k - $307k
...vehicle SoCs.In addition, you will optimize ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure... ...in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient...Full timeTemporary workRelocation package$216k - $293k
...leader. You will collaborate closely with clients and industry partners to identify opportunities. Additionally, you will support and develop a go-to-market team, contribute to delivery through billable roles with specific utilization targets, and partner with leadership...Temporary workLocal area$96.8k - $306.4k
...networking and routing, and a passion for building large-scale distributed systems to join our technical leadership team.Our team develops the core infrastructure that powers OCI Networking's data plane, solving complex challenges in routing, workload placement, and fleet...Temporary workFlexible hours- ...tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production... ...governed, and reproducible AI workflows.2+ years of experience developing APIs, integration services, automation workflows, or platform services...Full timeWork experience placementLive inWork at officeLocal area
$151.8k - $332.2k
...improve people’s work productivity.Responsibilities:Designing and developing scalable AI infrastructure solutions for training and deploying... ...AI environments using Docker and KubernetesOptimizing CUDA kernels for maximum GPU utilization and performanceDeveloping platform...Full timeWork at officeRemote work
