Data Scientist - PU Framework and Feature Store
CoAdvantage
About CoAd:
CoAd helps businesses navigate the complexities of workforce management through a combination of technology, expertise, and human support. Serving more than 17,000 clients nationwide, CoAd delivers payroll, HR, benefits administration, compliance support, workforce technology, and PEO services that help organizations simplify operations, support their employees, and drive growth.
Built on the belief that workforce solutions should be integrated, intuitive, and people-centered, CoAd empowers employers to manage their workforce with greater confidence while adapting to the evolving needs of today’s workplace.
Position Summary:
CoAdvantage is building two interlocking analytic substrates: the productivity-unit (PU) framework that measures realized τ across operational functions, and the analytics feature store that serves production ML models across the company. Both Data Scientist roles work across both substrates. This role reports to the AI Experimentation Lead and works in close coordination with the data team, the Staff MLOps Engineer, and the Principal AI Architect.
This is a hands-on engineering-grade data science role. The Data Scientist writes production-grade Python, owns model code from notebook to production, and uses AI-assisted coding tools as a daily driver. The role is not deck-only. The two Data Scientists collectively own the PU measurement layer and the model portfolio served from the feature store. Both roles are accountable for PU baselines and causal estimators on the measurement side, and for feature definitions and production models on the feature store side. Workload is allocated by the AI Experimentation Lead based on backlog priority and candidate strengths; neither role is scoped to a single domain. Active workstreams span MLR pricing redistribution, propensity to renew, churn, contact volume forecasting, staffing demand, payroll exception rates, and PU baselines for in-scope function.
Core Responsibilities:
- Model development:
Build production ML models that serve from the analytics feature store. This includes problem framing with stakeholders, feature engineering, model selection, training, validation, calibration, and packaging for production. Models are owned end-to-end; there is no separate "ML engineer" handoff. - Feature definitions in the feature store:
Author feature definitions, including source contracts, transformation logic, freshness requirements, and lineage metadata. The Data Scientist is accountable for the quality and correctness of every feature they introduce into the store. - Data team collaboration on feasibility:
Work directly with the data team to validate feasibility of proposed features before they are committed to the backlog: source system availability, refresh cadence, data quality, governance and tenant-isolation constraints. The Data Scientist is expected to be in the data team's review channels and to push back on infeasible features early rather than late. - PU baselines and causal estimators:
Both Data Scientists own portions of the PU measurement substrate: defining the productivity unit for in-scope functions, building the pipelines that compute c_r (c-sub-r) from operational data, and authoring the causal estimators that underwrite τ measurement — difference-in-differences, synthetic control, propensity matching, interrupted time-series. The Data Scientist is the analyst behind several of these readouts and is expected to be the methodological author of record on at least one PU baseline. - Model monitoring and reconciliation:
Define and instrument the production monitoring for each model: drift, calibration, business-metric reconciliation. Carry the on-call rotation for model issues alongside MLOps. - Documentation and methodological transparency:
Every model the Data Scientist ships carries a model card: assumptions, training data window, identification strategy, known failure modes, and the reconciliation plan. The bar is reproducibility from underlying data.
How we work:
- AI-first coding - Claude Code, Copilot, or successor tools are the default development surface. Feature definitions, model code, evaluation scripts, and monitoring instrumentation are expected to be authored with agentic coding tools in the loop. Hand-coding without AI assistance is the exception, not the norm.
- Hands-on with production code - The Data Scientist owns production code, not only research notebooks. Pull requests, code review, CI checks, and on-call all apply.
- Pre-registered targets - No model goes to production without a written success criterion and a reconciliation plan. The AI Experimentation Lead signs both.
- Methodological transparency - Identification strategies and validation choices are documented in writing and defended in review. "It performed well in cross-validation" is not sufficient.
- You estimate - Every workstream returns with a timeline, a confidence interval, and the smallest version that could ship in two weeks.
Required qualifications :
- Four or more years of experience as a Data Scientist or Applied Scientist building production models, not only research prototypes.
- Strong Python and SQL. Comfortable authoring production-grade analysis code, model training pipelines, and feature transformations without an engineering intermediary.
- Direct experience with at least one feature store — Feast, Databricks Feature Store, Tecton, Vertex Feature Store, or an internal equivalent — including authoring feature definitions and managing freshness.
- Demonstrated experience taking at least two models to production and operating them through at least one retrain cycle.
- Hands-on experience with AI-assisted coding tools (Claude Code, Copilot, Cursor, or equivalent) as a daily driver, with code commits or repositories to demonstrate the practice.
- Working fluency with causal inference (difference-in-differences, synthetic control, propensity matching) AND at least one of: time-series forecasting, propensity modeling, constrained optimization. Both roles are expected to operate across PU measurement and feature-store modeling, so causal methods are not optional.
- Written communication skills sufficient to produce model cards and stakeholder-facing readouts.
Preferred qualifications :
- Prior experience in a PEO, HR outsourcing, insurance brokerage, BPO, or other labor-intensive services organization.
- Direct exposure to pricing, underwriting, churn, or operational forecasting use cases.
- Familiarity with cloud ML platforms (Azure ML, Databricks, Vertex AI).
- Experience working under data governance constraints typical of regulated multi-tenant environments (HIPAA, PII, tenant isolation).
What success looks like at 12 months:
- At least two production models owned end-to-end, with documented model cards and active monitoring.
- At least eight feature definitions contributed to the feature store, with the data team co-sign on lineage and freshness.
- At least one PU baseline (c_r) published with audit trail and methodological documentation.
- At least one causal readout co-authored with the AI Experimentation Lead that informed an executive-level tooling or pricing decision.
- Established working pattern with the data team — feasibility reviews routine rather than ad hoc.
EEO
CoAdvantage is committed to providing equal employment opportunities to all employees and applicants without regard to race, color, religion, national origin, ancestry, citizenship status, age, sex (including pregnancy, childbirth, breast feeding and pregnancy-related medical conditions), gender, gender identity or expression, sexual orientation, marital status, uniform service member and veteran status, disability, genetic information, or any other characteristic protected by applicable federal, state, or local laws and ordinances.
#LI-remote- ...Data Scientist The Data Scientist leads the team in developing sophisticated predictive models... ...can interact with large amounts of data stored in a Hadoop environment. Is capable... ...techniques including; variable selection, feature engineering, model generation, model...Suggested
- ...community of in-person work. About the Role As a Forward Deployed Data Engineer — SIEM/SOAR, you build the content that powers TENEX's... ...for diverse data source types Knowledge of MITRE ATT&CK framework and its application to detection content Experience with Python...Suggested
- ...Summary: CoAdvantage is adding a dedicated Data Scientist to conceive, prototype, and validate new... ...latitude to define the problem framing, feature strategy, and experimental design for... ...ML, Databricks, Vertex AI) and feature store concepts. Experience working under...SuggestedPermanent employmentContract workLocal areaRemote work
- ...large-scale LLMs, graph-based reasoning engines, and streaming feature pipelines that operate on billions of security events.... ...Graph structures and specifically graph databases. Orchestration Frameworks: Hands-on experience building agents, orchestration frameworks...SuggestedWork from home
- ...RAG-style workflows (feeding documents or data into LLMs in a controlled, auditable way... ...into scoped, deliverable AI-powered features and tools. Take ownership of projects... ...or business problem. Exposure to web frameworks (FastAPI, Flask, or a front-end framework...SuggestedWeekly payFull timeInternshipWork at office
$149k - $170k
...offer a video platform, cloud services, advertising solutions, and a non-custodial cryptocurrency wallet. Rumble is seeking a Senior Data Engineer to design, build, and operate the data platforms and backend systems that support large-scale product, analytics, and...$80k
...Senior Data Analytics Engineer Sarasota County Government has an excellent opportunity to be a Senior Data Analytics Engineer at... ...relational data including data pipelines, system interfaces, and data stores. Strong scripting skills in Python,.NET, PowerShell, and MS...Temporary workMonday to Friday2 days per week- ...and SIGINT processing needs, including working with time-series data. Develop specialized models for event characterization, pattern... ...processing systems for dynamic environments. Discover features and infer system states from underlying data streams to support...Hourly payFull timeWeekend work
$132.9k - $174.45k
At Jacobs, we’re challenging today to reinvent tomorrow by solving the world’s most critical problems for thriving cities, resilient environments, mission‑critical outcomes, operational advancement, scientific discovery and cutting‑edge manufacturing, turning abstract ...Full timeWork at officeLocal areaRemote work$135k
...SIGINT processing needs. This includes working with time-series data and developing models for event characterization, pattern... ...adaptive processing systems for dynamic environments, and discovering features and inferring system states from the underlying data streams....Relocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Scientist - PU Framework and Feature Store. Be the first to apply!


