Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
$200k
...Lead Data Engineer - Remote 100% Remote (U.S. Based) Up to $200,000 Base Salary U.S. Citizens & Green Card Holders Only I'm partnering with a rapidly growing healthcare technology company that is looking to hire a Lead Data Engineer to help drive the evolution...SuggestedRemote work$100k - $120k
...Data Engineering - DevOps Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in the enterprise. Fractal is building a world where individual choices, freedom, and diversity are the greatest assets. An ecosystem...SuggestedHourly payFull timeLocal areaRemote work- ...Location: Remote We are seeking an experienced Data Engineer to support challenging data engineering, data management, and analytics initiatives. In this role, you will develop and maintain data pipelines and data structures that enable advanced analytics and data...SuggestedTemporary workLocal areaRemote work
- ...solutions that enable seamless communication and collaboration across the supply chain. DESCRIPTION: Transflo is seeking a Senior Data Engineer to architect and own our enterprise data platform — from raw ingestion through curated, analytics-ready data products. You will...SuggestedRemote work
- ...Responsibilities: Build and operate production-grade batch and streaming data pipelines across SQL Server, cloud applications, device/event... ...on-call activities. Work closely with Product, Application Engineering, Quality Assurance, DevOps/Site Reliability Engineering,...Suggested
- ...customers' business challenges, Take2 will work as a partner to best resolve client needs. Take2 is hiring an AWS Lakehouse Data Engineer who is eligible to be sponsored for a Public Trust Clearance. This position is Remote, but it will require you to work East Coast...Remote work
- ...Data Engineer IICrane Aerospace and Electronics has an exciting opportunity for a Data Engineer II at our Lynnwood, WA location. This is an on-site position.Crane Aerospace & Electronics supplies critical systems and components to the aerospace and defense markets. You...For contractorsWork at office
$140k - $180k
...Job title: Energy Storage Sales Engineer – AI Data Centers Location : Seattle, WA United States - Remote Pay Rate : $140,000 to $180,000/year (Upon negotiation) Schedule : Full-time Duration : Permanent Requirements: Experience in at least one of...Permanent employmentFull timeWork experience placementRemote work$210k - $250k
...What You Will Be Doing: Helion is hiring a Senior Applied Plasma Data Scientist specializing in synthetic diagnostics to improve how... ...for research stakeholdersRequired Skills:PhD in Physics, Engineering, Applied Math, Computational Science, Machine Learning, or related...Temporary workWork at office- ...strengthening the resilience of the U.S. financial system. Data Scientists at Zest AI use the power of machine learning (ML)... ...explainability, and regulatory constraints Collaborate with ML engineers, product leaders, and domain experts to bring models from...Work at office
- ...Data Scientist Remote, Anywhere in the Continental US Due to the nature of the role, this position requires that you are a U... ...technical and non-technical audiences. Collaborate with policy, engineering, and operations teams to ensure analytics solutions align with...Full timeImmediate startRemote work
- ...DIRECTLY ON OUR W2 (EAD, OPT, USC, GC, H4) NO THIRD PARTY! Data Scientist We are seeking a Data Scientist to join a growing... ...degree in Data Science, Statistics, Mathematics, Computer Science, Engineering, or a related field. - FROM THE USA ~3+ years of experience...
- ...Great Place to Work® certification year after year. Principal Data Scientist Job requirements ~ Experience Range: 15 - 18... ...PySpark, and R, ensuring efficient data processing and feature engineering Develop, validate, and maintain probabilistic graph models and...
$42.55 - $66.06 per hour
Positions at this level serve as Epic application subject matter experts for Cadence and Prelude, applying advanced knowledge of healthcare operations, clinical/business workflows, and Epic system functionality to support a large, integrated healthcare organization. Analysts...Minimum wageLocal areaRemote workShift work- ...infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent. About the Role: We're looking for a high-taste engineer with a strong eye for quality — someone who'...For contractorsFreelanceImmediate startRemote work
- ...supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by...For contractorsRemote workFlexible hours
- ...infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent. About the Role: We're looking for a high-taste engineer with a strong eye for quality — someone who'...For contractorsFreelanceImmediate startRemote work
- ...in technology. Job Description Analyze and organize raw data Build data systems and pipelines Conduct complex data analysis... ...algorithms, and operating systems) 0-1+ years of data engineering experience Benefits Job technical support E-verified...Full time
$150k - $180k
...as leading organizations at . About the Job The Cloud Data Architect is a customer-facing mid-senior level (depending on experience... ..., and/or Terraform preferred) and daily work alongside data engineering and/or software engineering teams. Experience in developing...Full timeFor contractorsLocal areaRemote workRelocation$35 per hour
Seasoned IT Analyst / Service Desk Analyst with VERY strong customer service and communication (both verbal and written) skills. Must be detail-oriented and able to lead the development of, or follow established processes and procedures, to execute the following responsibilities...Hourly payWork at office- ...deliver them in as little as one hour. We're hiring a Senior Machine Learning Engineer for our Personalization Platform team within the Membership organization. You'll partner closely with Data Scientists to design, deploy, and scale personalized recommendation and...Work at officeFlexible hours
- ...Position Overview A large grocery retailer is seeking a Lead AI Engineer / Agentic Commerce Lead to drive the architecture, development, and delivery of customer-facing AI agent experiences. This is a highly technical leadership role designed for an experienced engineer...Local area
- ...AI Engineer Location: Remote, Nationwide Our client is building the next generation of intelligent software, where AI is expected to do more than generate a compelling response. The goal is to create dependable AI experiences that understand context, take action...Remote work
- ..., with new products for governments in R&D. We are hiring an engineer to work alongside our founders and team. You will support rapidly... ...everyday technical work a small team needs: debugging, scripts, data, integrations. Who you are - Curious and adaptable. You...Live inRemote work
- Company DescriptionSA Technologies Inc. (Please visit for more details about our organization.) is a market leader and one of the fastest growing IT consulting firms with operations in US, Canada, Mexico & India. SAT is an Oracle Gold Partner, SAP Services Partner & IBM...
- Earn at Home by Taking Polls - Data Entry Clerk - Customer Service Rep - Work at Home & Part TimeWe are looking for people nationwide to participate in polls - Apply ASAP!We offer you the opportunity to earn extra income from home (teleworking) and also to decide your...Extra incomePart timeImmediate startWork from home
$41.35 per hour
...solves repair problems by studying drawings, wiring diagrams and schematics, and technical publications; uses automated maintenance data systems to monitor maintenance trends, analyze equipment requirements, maintain equipment records, and document maintenance actions,...Contract workLocal area- Job Title: Backbase Developer Location: Remote Job description - Must Have Technical/Functional Skills • Experience with Backbase. • Strong programming skills in .NET or Java or React JS • Hands-on experience with API development and integration. • Familiarity...Remote work
- ...Platform Engineer NeuBird AI is scaling rapidly and we need a platform engineer who can build the internal tools and infrastructure that make our engineering teams more productive. You'll create self-service platforms, streamline developer workflows, and eliminate friction...Flexible hours
- ...business challenges? Join Korry as our Enterprise AI & Automation Engineer and lead the development of innovative AI solutions that... ...Server, Azure, Microsoft Graph, and third-party SaaS platforms.Data & AnalyticsDevelop Power BI dashboards and executive reporting.Support...Work at officeNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineering intern summer Everett, WA
- ai data Everett, WA
- data loss prevention engineer Everett, WA
- provider data management Everett, WA
- data cabling Everett, WA
- health data Everett, WA
- test data management Everett, WA
- data cabling installation Everett, WA
- data internship Everett, WA
- data modeling Everett, WA



