Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
- ...Data Engineer & Analyst — Federal Analytics Consulting Location: Washington, DC Metro Area — Hybrid (3 days remote / 2 days on-site) Employment Type: Contract / 1099 / Contract-to-Hire, converting to full-time as you prove yourself Eligibility: U.S. Citizenship or Permanent...SuggestedPermanent employmentFull timeContract workImmediate startRemote workRelocation
- ...from idea to viable formulation faster by unifying fragmented R&D data: ingredients, specs, cost, nutrition, processing constraints and... ...THE ROLE We are looking for a motivated, hands-on Data Engineer with 3+ years of experience to join our team. You will build...SuggestedFull time
- ...Role: Data Engineer - Databricks Location: Washington, DC Metro Area Duration: Long-Term (W2) USC (Does not require Sponsorship) Experienced data engineers to join a large-scale Databricks platform build supporting federal government programs in the Washington...SuggestedFull timeRemote workRelocation
- ...Data Engineer (Databricks) Location: Hybrid (Washington D.C) 2-3 times a week Clearance: Active Secret Employment Type: Full Time Company Description Big Impact Tech (BIT) is a Small Business providing IT and business management consulting to federal and...SuggestedFull time
- ...Data Engineer Every day, healthcare providers across the country navigate systems that are supposed to make care better, faster, and more affordable - and too often don't. Behind every claim, every quality measure, every interoperability standard is a doctor trying to...SuggestedFull time
$137.7k - $229.5k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ..., and fuel growth through innovation.Work You’ll DoAs a Lead Data Engineer II on the team, you will be responsible for:Support...Local area$128.1k - $166.2k
...your potential, but also to contribute to our clients’ success.Data EngineerAbout the DepartmentThe Client Data Services (CDS) team... ...the firm.CDS operates as a cross-functional team combining data engineers, software engineers, and analysts who partner closely with...Work experience placementLocal area3 days per week$105.4k - $207.8k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ..., and fuel growth through innovation.Work you'll do As a Data Engineer III on the AI & Data team, you will be responsible for...Local area$100.79k - $160.26k
...DLA Piper is a place you can engage in meaningful work and grow your career. Let’s see what we can achieve. Together.SummaryThe Data Engineer, Solutions & Data role designs, builds, and operates data pipelines and data integration processes that translate raw data into...Full timeWork at officeRemote workVisa sponsorship$84.4k - $140.6k
Position Summary Deloitte is seeking a Data Engineer II to support the design, development, and delivery of data solutions that enable analytics, reporting, and business decision-making. This role will focus on building scalable data pipelines, improving data quality...- ...This role is primarily focused on data engineering, transformation frameworks, orchestration, and system integrations. While the team builds applications on top of their platform, dedicated application engineering resources lead full-stack UI development. This position...Full timeFlexible hours
- ...Location: Remote We are seeking an experienced Data Engineer to support challenging data engineering, data management, and analytics initiatives. In this role, you will develop and maintain data pipelines and data structures that enable advanced analytics and data...Temporary workLocal areaRemote work
- ...Senior Data Engineer With Databricks Remote, Anywhere in the Continental US Due to the nature of the role, this position requires U.S. Citizenship. Data Surge is disrupting the services industry with cutting-edge technology that brings together the very best...Full timeImmediate startRemote work
- ...Sand Sand Technologies is a global Physical AI company using data and AI to make critical industries work better. We partner with... ...the US, we operate across the full stack - from research and engineering to deployment and capability building. Our mission is simple...
- ...Data Engineer IIThe Data Engineer II will be responsible for designing, building, and maintaining the bank's data infrastructure and systems. This role involves working with cross-functional teams to ensure efficient data integration, transformation, and storage. The Data...Work at office
$100k - $120k
...Data Engineer IIInVita Healthcare Technologies is a leading software provider for complex medical, forensics, and community care environments. We build specialized, highly configurable, and integrated systems that support hospitals, blood centers, donation organizations...Work at officeLocal areaMonday to Friday3 days per week$100k - $120k
...Data Engineering - DevOps Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in the enterprise. Fractal is building a world where individual choices, freedom, and diversity are the greatest assets. An ecosystem...Hourly payFull timeLocal areaRemote work- ...Data EngineerThe Data Office team at T. Rowe Price is playing a key role in helping build the future of financial services, working... ...people invest. We are seeking a highly skilled and experienced Data Engineer to join our team and play a critical role in building and...Work at office
$121.6k - $163.8k
...staff experience through career development and educational opportunities. Position Overview Index Analytics is seeking a Data Engineer to support Government clients to design, build, and optimize scalable data pipelines and cloud-based solutions. The Data Engineer...Local area$121.6k - $163.8k
...Data EngineerFully RemoteOverviewSalary Range $121,600.00 - $163,800.00 Salary/year Level Senior Education Level 4 Year DegreeDescriptionCompany... ....Position OverviewIndex Analytics is seeking a Data Engineer to support Government clients to design, build, and optimize scalable...- ...Skills ~12+ years of solid hands-on experience in Python, AWS Data Services, DBT, Apache Airflow (on Astronomer platform), SQL and... ..., etc. Guide, coach, and upskill junior and mid-level data engineers on best practices, coding standards, and modern data patterns...
$100k - $140k
...Data Engineer Opportunity At ResilienceStep into a critical technical role driving the core data architecture that powers our category-defining platform. As a Data Engineer at Resilience, you won't just maintain existing pipelines; you will take full ownership of building...Full timeRemote work- ...Position Purpose: We are seeking a Data Engineer to support the development of cloud-native data solutions that improve operational efficiency, support regulatory reporting, and drive actionable insight for our commercial and government healthcare clients. This role...
- ...from idea to viable formulation faster by unifying fragmented R&D data: ingredients, specs, cost, nutrition, processing constraints,... ...looking for a curious, motivated, and detail-oriented Junior Data Engineer to join our team. This role is ideal for someone with strong...Full time
$87.64k - $153.55k
We are seeking a Data Engineer I to assist with the creation and maintenance of complex data pipelines from raw acquisition of data to visualization. The Data Engineer I will support the design, production, and maintenance of software infrastructure to automatically extract...Full timeFor contractors$102k - $136k
...Senior Data EngineerThe Senior Data Engineer is a highly hands-on individual contributor responsible for building, operating, and improving Sinclair's enterprise data platform. This is a Snowflake-first engineering role: most of the work is performed within Snowflake and...Full timeWork experience placementH1bLocal areaFlexible hours- ...Quantiphi: Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services... ...by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our...Full timeRemote work
- ...Responsibilities: Build and operate production-grade batch and streaming data pipelines across SQL Server, cloud applications, device/event... ...on-call activities. Work closely with Product, Application Engineering, Quality Assurance, DevOps/Site Reliability Engineering,...
- ...Mission-driven digital transformation firm is looking to add a Senior Data Engineer to its growing engineering team in Baltimore, Maryland (open to full-remote). This is an opportunity to grow into a leadership role whilst building modern cloud data platforms and scalable...Remote work
- ...solutions that enable seamless communication and collaboration across the supply chain. DESCRIPTION: Transflo is seeking a Senior Data Engineer to architect and own our enterprise data platform — from raw ingestion through curated, analytics-ready data products. You will...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineer machine learning Baltimore, MD
- aws data engineer Baltimore, MD
- sr data engineer Baltimore, MD
- big data cloud engineer Baltimore, MD
- finance data engineer Baltimore, MD
- entry level data engineer Baltimore, MD
- sr information security engineer Baltimore, MD
- data center engineer Baltimore, MD
- senior data integration developer Baltimore, MD
- senior cloud data engineer Baltimore, MD



