Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
- Find a JobJoin IDR Internal TeamStaffing ServicesConsultantsAbout UsBlogSuggested
- ...Business Oriented IT environment with rich involvement in technology innovation, ERP and CRM counselling, Product Engineering, Business Intelligence, Data Management, SOA, BPM, Data Warehousing, SharePoint Consulting and IT Infrastructure. Our other offerings include modified...SuggestedContract workLocal areaWorldwide
$200k
...Lead Data Engineer - Remote 100% Remote (U.S. Based) Up to $200,000 Base Salary U.S. Citizens & Green Card Holders Only I'm partnering with a rapidly growing healthcare technology company that is looking to hire a Lead Data Engineer to help drive the evolution...SuggestedRemote work$100k - $120k
...Data Engineering - DevOps Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in the enterprise. Fractal is building a world where individual choices, freedom, and diversity are the greatest assets. An ecosystem...SuggestedHourly payFull timeLocal areaRemote work$102k
...modifies, configures, debugs and evaluates jobs for extracting data from various sources, implements transformation logic, and stores... ...discipline or equivalent experience Experience with data engineering/ETL ecosystems (e.g., Informatica, SAP BODS, OBIEE), 3 yrs Desired...SuggestedWork at officeRemote work2 days per week1 day per week- ...We are seeking a Data Engineer with expertise in Azure Databricks to support the modernization of enterprise data platforms. The ideal candidate will have experience migrating Oracle PL/SQL-based ETL processes to Databricks, developing scalable cloud-based data pipelines...
$110.93k - $166.4k
...The Data Engineer designs, builds, and supports the data pipelines, integrations, and curated datasets that power analytics, reporting, and AI initiatives for PBK's Architecture vertical and broader AEC operations. This role owns integrations across enterprise platforms...- ...Location: Remote We are seeking an experienced Data Engineer to support challenging data engineering, data management, and analytics initiatives. In this role, you will develop and maintain data pipelines and data structures that enable advanced analytics and data...Temporary workLocal areaRemote work
$170k - $180k
...Data Engineer OpportunityPremier Nutrition Company (PNC) is one of the fastest-growing companies in the proactive wellness space, showing clear leadership in the category of protein shakes and powders. We make the brands Premier Protein and Dymatize, and are part of our...Temporary workWork at officeDay shift- ...Senior Data EngineerThe Data Services team is responsible for technical design and end-to-end delivery of complex data-driven solutions and data products for the enterprise. The Senior Data Engineer will report to the Sr Manager, Data Solutions / Manager. This role is...Full timePart timeWork at officeLocal areaWork from homeHome office2 days per week
- ...Oakland , CA / Seattle , WA / Providence , RI ... View All The Data Services team is responsible for technical design and end-to-end... ...driven solutions and data products for the enterprise. The Data Engineer,Senior will report to the Sr Manager, Data Solutions / Manager....Local area
$180k - $220k
...Senior Data EngineerBerkeley, CAAt Aircapture we're creating technology to solve what we believe to be our lifetime's most pressing... ...our Direct Air Capture (DAC) systems into the insights that our engineers and scientists rely upon. Partnering with our test and...$116k - $145k
...Senior Data Engineer role on the Anthro Data Team to build and scale the data backbone across pipelines, modeling, AWS infrastructure, monitoring, and data quality. Responsibilities Design and build production-grade ELT pipelines for lab and manufacturing...Flexible hours- ...Job Title: Senior Data Engineer Location: Berkeley, CA (Onsite – 5 Days/Week) Employment Type: Contract/Full-Time About the Role We are seeking a Senior Data Engineer to join an innovative company in the industrial technology space. This is an opportunity to...Full timeContract work
- ...Senior Data EngineerWorldly is the world's most comprehensive impact intelligence platform — delivering real data to businesses on impacts... ...mission-aligned investors.Worldly is hiring a hands-on data engineer with a passion for sustainability to join our dynamic team. You...Work at officeRemote workWork from homeFlexible hoursShift work
$180k - $220k
...our goals. Thank you for considering us. You will shape the data foundation that powers the technology of our lifetime, owning the... ...our Direct Air Capture (DAC) systems into the insights our engineers and scientists rely upon. Partnering with cross-functional stakeholders...- ...Quantiphi: Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services... ...by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our...Full timeRemote work
- ...solutions that enable seamless communication and collaboration across the supply chain. DESCRIPTION: Transflo is seeking a Senior Data Engineer to architect and own our enterprise data platform — from raw ingestion through curated, analytics-ready data products. You will...Remote work
$150k - $180k
...Lead Control Engineer // Data Center Developer Location: Austin, Texas (On-site)Relocation Needed Compensation: $150,000–$180,000 base About the Company A leading developer of hyperscale data centers, delivering scalable, efficient, and sustainable solutions...RelocationFlexible hours$75 - $81 per hour
...IDR is seeking a Geospatial Data Engineer to join one of our top clients for a remote contract opportunity (California preferred) . This role is ideal for a senior Data Engineer with deep experience building and running geospatial/raster data pipelines natively on AWS...Contract workTemporary workRemote work- National Science Foundation funded PhD/Postdoc Lab at UC Berkeley is seeking researchers to advance a declarative data management system for semantic multimodal workflows. The project focuses on documents and videos, offering a high-level DSL and low-code interface to...
$140k - $180k
...Job Description Job Description Job Title: Senior Data Engineer Industry: Healthcare Location: Emeryville, CA Assignment Type: Full time, direct hire Expected Salary: $140-180k Work Schedule: Hybrid, 1 day a week Benefits: This position is eligible...Full timeLocal areaShift work1 day per week$15k
...become a multibillion-dollar asset manager, and we have ambitious goals for the future.Your TeamAs a Senior or Staff Software Engineer on our Data Engineering team, you will contribute to scaling and advancing our entire data operation. This includes procurement and...Local area$185k - $223k
...Staff Data EngineerAt Copper, we're reinventing home appliances for an electrified future. Our flagship Charlie range pairs high-performance... ...workflows — e.g., product triage, anomaly flagging — across engineering, support, and product.What you'll bring8+ years building...Full timeWork at officeRemote workFlexible hours- ...Staff Data EngineerForm Energy is hiring a Staff Data Engineer to build the data systems behind our deployed energy storage projects, turning real-time field telemetry into reliable, trustworthy data the rest of the company can build on. As a hands-on senior individual...Immediate startRelocation package
- ...become a multibillion‑dollar asset manager, and we have ambitious goals for the future. Role Overview As a Staff Software Engineer on our Data Engineering team, you will contribute to scaling and advancing our entire data operation. This includes procurement and...
$50 per hour
...Unit: Enterprise Business ServicesStandard Job DescriptionJoin Lockheed Martin’s Data & AI Enablement organization and accelerate the company’s AI transformation. As a lead data engineer, you will design, build, and productize data assets that power AI agents, analytics...Full timeTemporary workWork experience placementCasual workFlexible hours- Trajectory Labs, PBC is hiring a Tech Lead to build a software factory for safety data used to train and evaluate frontier AI models. You will own complex technical problems from experiments to production, partnering with researchers to create training datasets, environments...Remote work
- ...customers' business challenges, Take2 will work as a partner to best resolve client needs. Take2 is hiring an AWS Lakehouse Data Engineer who is eligible to be sponsored for a Public Trust Clearance. This position is Remote, but it will require you to work East Coast...Remote work
$155k
...team within the overall Electric Transmission and Distribution Engineering organization is responsible for planning, organizing, and managing... ...and initiatives. Within this department the Reliability Data team is on point for a key role is developing and curating all...Work at officeRemote work2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineer machine learning Oakland, CA
- finance data engineer Oakland, CA
- data center engineer Oakland, CA
- senior cloud data engineer Oakland, CA
- data engineer Oakland, CA
- data engineer analytics Oakland, CA
- senior data center engineer Oakland, CA
- data science developer Oakland, CA
- junior data engineer remote Oakland, CA
- data developer Oakland, CA



