Senior AI Support Engineer: Scale Reliability & Incident Response
United States Digital Space LLC
The Senior Support Engineer role is remote-based in Toronto, Canada, focusing on collaborating with strategic enterprise accounts and product teams to resolve the most challenging issues on our API platform. You will help design and run monitoring, alerting, and incident response processes to ensure reliability at scale. You will work closely with engineering and infrastructure teams, tackling high-difficulty problems in a low-volume environment and providing coverage during peak periods as #J-18808-Ljbffr United States Digital Space LLC
- ...Robinhood is hiring a Senior Engineer who will join the newly formed reliability team, responsible for coordinating production incident responses. This role emphasizes leadership during incidents... ...to help shape incident management processes at scale. #J-18808-LjbffrSenior
- ...Senior Software Engineer –Gen AI Platforms Job Req Id: 26961645 Location(s):... ...edge products at planetary scale. We are solely focused... ...build high-quality, highly reliable, and secure distributed systems... ...to bug bounty, responsible disclosure programs, and AI...SeniorFull timeWork at office
$120k - $165k
...a Software Engineering position at... ...job family responsible for developing... ...solutions that support business... ...Analytics teams, Senior management,... ...-based AI applications... ...production, reliability, and evolution... ...services for large-scale enterprise... ..., incident management,...SeniorTemporary work$150k - $160k
...productionization of AI capabilities across the... ...-powered products at scale and now wants to... ...this role.THE ROLEAs a Senior AI Engineer (Full-Stack / Applications... ...quality, latency, reliability and cost.Collaborate... ...in on-call rotations, incident response and post-mortems for...SeniorFull timeContract workTemporary workWork at officeLocal areaRemote work- the company, a AI research and deployment firm, is seeking... ...seasoned User Operations engineer to help customers adopt AI at scale. You will troubleshoot... ...own complex production incidents in a fast-moving environment... ...scalable, AI‑driven support workflows. Remote work from...SeniorRemote job
$123k - $215.25k
...back the broader engineering community... ...and services in support of Amex’s customers... ...Leader for its Incident Response team. This is a senior-level, hands-on... ...adoption of AI-enabled security... ...sophisticated threats at scale. This role... ...accuracy, reliability, explainability...SeniorShift work$113.1k - $232.3k
...Lead Applied AI Site Reliability Engineer II Role Overview... ...resilient systems at scale. The ideal... ...confidence. Key Responsibilities: Outcome-Driven... ...environments you support to meet their SLOs... ...budget, and track incident trends and toil... ...level employees to senior leaders, we...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week$84.9k - $188.1k
Senior VMware Virtualization Support Engineer Position Description CGI... ...platform maintenance, incident management, and... ...engineer is responsible for maintaining... ...automation and AI-enabled... ...improve platform reliability and support efficiency... ...large-scale environments using...SeniorLocal areaDay shift- ...BlackLine is seeking a Senior AI/ML Engineer to design, build, and optimize data pipelines powering... ..., and ensure data accuracy at scale across cloud environments. You’ll guide... ...define governance and observability practices for reliable AI systems. #J-18808-LjbffrSenior
$158.5k - $172k
...The OpportunityAs a Senior Engineer on the Runtime Automation... .... Our team is responsible for managing our centralized... ...driving continuous reliability, deep system... ...continuous knowledge sharing.Scale Infrastructure as Code... ...-call rotation as an incident responder, leading...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week$141k - $216.6k
...infrastructure with an AI intelligence... ...of emergency response and building a safer... ...OverviewAs a Site Reliability Engineer, you'll own the... ...operational practices that support our applications... ..., alerting, and incident response across... ...continues to scale.What Make You a...SeniorWork experience placementWork at office$104k - $124.8k
...experience building and supporting large-scale Python... ...integration, prompt engineering, fine-tuning or... .... Responsibilities: We design and... ...implement agentic AI systems, including... ...observability, incident management, and... ...functions tied to reliability, risk reduction,...SeniorHourly payFull timeRemote work$139k - $257.55k
...seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native... ...ownership. On-call and incident response come with the territory,... ..., and operate large-scale, distributed, fault-tolerant...SeniorFull timeTemporary workLocal areaRemote workWorldwide$125k - $145k
...trusted trade. Our AI-powered product... ...and passionate Senior IT Support Engineer to join our... ...shaping how a fast-scaling, AI-forward company... ...team on incidents, access, and hardening... ...conditions, and ensure reliability over time.... .... Our responsibility extends beyond individual...SeniorPermanent employmentFull timeTemporary workWork experience placementWork at officeFlexible hours$167.7k - $245.2k
...role, you will be responsible for maintaining services... ...t own. Powered by AI and an unmatched... ...deploy at scale while also delivering... ...looking for talented engineers with a software or... ...to ensure the reliability, performance and security... ...our 24x7 incident response and on-call...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$150.9k - $226.3k
jobr.pro is seeking an Incident Response Technical Program Manager in New York, NY. This senior role involves leading security and reliability incidents at Harvey, a fast-growing AI company. Responsibilities include end-to-end incident coordination, designing management...Senior$182.3k - $220k
...systems that operate at scale, with an opportunity... .... We're hiring a Senior AI Engineer to join a new team we... ...that make LLMs reliable enough for real-world... ...resumes, and assessing responses or job-related qualifications... ...tools are used to support our recruitment team...SeniorFull timeLocal area$161.8k - $184.6k
Senior AI Engineer Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years... ...develop, test, deploy, and support AI software components... ..., throughput - of large scale production AI systems....SeniorFull timePart timeLocal area$144k - $236k
...business needs of the team. Responsibilities: AI is at the core of how... ...trust platforms. As a Senior AI Software Engineer you will own end-to-end machine... ...production at LinkedIn scale. You won't just train... .... You operate systems reliably as on-call and root-cause...SeniorFull timeFor contractorsWork at officeImmediate startFlexible hours- As an AI Platform Engineer for AI & Emerging Tech... ...on technical support to AI... ...users understand responsible AI usage, credit... ...to scale support.... ..., usage, and incident data to drive... ...improvement, reliability, and platform... ...Seniority Senior...Senior
$160k - $220k
...so our infrastructure must support rapid scalability on holidays... ...our small infrastructure engineering team in our NYC and SF offices... ...a Rails API. These systems scale reliably while supporting rapidly... ...monitoring, alerting, and incident response Reduce application and database...SeniorFull timePart timeFlexible hours$198k - $264k
...Essential Cloud for AI. Built for... ...to build and scale AI with... ...TeamThe Technical Support Engineering - Infrastructure... ...the RoleAs the Senior Manager of... ...support scale reliably. You'll lead with... ...including roles and responsibilities, escalation... ...critical incidents and resolve conflicts...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$229.9k - $262.4k
Senior Lead AI Engineer Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital... ...develop, test, deploy, and support AI software components... ...latency, throughput — of large scale production AI systems....SeniorFull timePart timeLocal area$123k - $215.25k
...benefitsJob Function: Engineering &... ...a commitment to supporting the broader engineering... ...American Express, AI is reshaping the... ...capabilities with the reliability, security,... ...American Express scale.THE ROLEAs a Senior AI Engineer -... ...hands-on engineer responsible for designing, building...SeniorWorldwideVisa sponsorship$204k - $255k
...s Unstructured AI TeamWork at the... ...forefront of context engineering - shaping how... ...results at scale.Own end-to-end... ...progress as a team.Senior AI Engineer at Collibra are responsible forShipping... ...deliver accurate, reliable results.... ...flexibility in mind to support you and your...SeniorWork experience placementWork at officeFlexible hours2 days per week$144k - $236k
...business needs of the team.Responsibilities: AI is at the core of how... ...and trust platforms. As a Senior AI Software Engineer you will own end-to-end machine... ...production at LinkedIn scale. You won't just train... ...data.You operate systems reliably as on-call and root-cause...SeniorFor contractorsWork at officeImmediate startFlexible hours- ...Merciv Merciv is an AI and data... ...around a strong engineering culture, and are... ...a DRI (Directly Responsible Individual) model... ...The Role As a Senior Backend Engineer,... ...ensure our platform scales reliably for Fortune 500 clients... ...AI platform, supporting real-time decision...SeniorFull timeImmediate startHome officeFlexible hours
$104.9k - $174.7k
...or infrastructure engineering roles We need... ...need experience supporting production environments through incident response and on-call rotations... ...supporting large-scale, highly available... ...teams to improve reliability, performance, and... ...We are hiring a Senior Site Reliability...SeniorFull timeWork at officeRemote work$500 per month
...diverse group of experienced engineers, traders, and brokerage... ...apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help... ...day-to-day - oncall, incident response, postmortems, and the follow... ...ownership, connection pooling at scale, or change-data-capture pipelines...SeniorHome office$130k - $147k
...decision. Our AI-first... ...are seeking a Senior ML/AI Platform Engineer to help build... ...engineering, and site reliability, with a focus... ...systems at scale. This is a... ...You will be responsible for designing... ...capabilities. Support governance,... ...improvements,incident response, root...SeniorPart timeWork at officeRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Support Engineer: Scale Reliability & Incident Response. Be the first to apply!
- ai ml engineer New York, NY
- machine learning ai engineer New York, NY
- ai developer New York, NY
- ai engineer New York, NY
- ai research engineer New York, NY
- ai prompt engineer New York, NY
- senior ai engineer New York, NY
- ai engineer remote New York, NY
- IT developer New York, NY
- technical support engineer New York, NY



