Data Scientist
Sunset
Sunset Data Scientist Role
Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. That creates a difficult measurement problem. A system can improve aggregate F1 while missing a high-risk slice, remove more sensitive information while also destroying useful context, or pass one stage while defects escape somewhere else in the pipeline.
As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers.
This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships.
You will work closely with Machine Learning, Product Engineering, Data Engineering, Security, Quality, domain experts, and the team making delivery decisions. Machine Learning Engineers own changing model behavior. You own the credibility of the evidence used to decide whether a model, pipeline, or delivery change actually made the data safer or more useful.
Questions You Might Answer
- Did a higher NER or entity-resolution score actually reduce sensitive misses across the messages, documents, tables, and providers that matter?
- Is a new model finding more sensitive information, or simply removing more of the useful structure our customers need?
- Can we trust a golden dataset, a human review process, or an LLM judge enough to use it for a release decision?
- Which customer, modality, entity, language, or format slices are hidden by a strong aggregate result?
- Where did a quality loss enter between source data, processing, de-identification, review, and delivery?
- What is the smallest credible experiment that would tell us whether to ship, revise, or stop a change?
What You'll Do
- Define what high-quality and safe-to-deliver data mean across de-identification, structure preservation, semantic coherence, and customer utility
- Design representative samples and build golden, adversarial, replay, and production-like corpora with explicit provenance, labeling policy, agreement, adjudication, and versioning
- Turn ambiguous concepts such as "useful," "clean," or "safe" into measurable claims with known uncertainty and clear decision consequences
- Evaluate detectors, models, prompts, judges, thresholds, review workflows, and pipeline changes using comparisons that can support a real decision
- Break aggregate results into the modalities, providers, entity classes, customer contexts, languages, formats, and risk tiers that reveal consequential failures
- Connect local measures to escaped sensitive information, avoidable over-redaction, preserved data utility, review burden, rework, and delivery acceptance
- Build reproducible analysis, evaluation pipelines, and high-fidelity environments using Python, SQL, synthetic data, historical replay, seeded failures, and programmatic verifiers
- Establish holdout and evaluation practices that keep the evidence trustworthy while model and product teams iterate quickly
- Use modern AI tools deeply for analysis, corpus development, coding, review, and hypothesis generation while independently verifying their output
What Success Looks Like
- The team has a decision-grade baseline for a priority Clean Data quality claim and trusts it enough to use in model, pipeline, release, and delivery decisions
- Improvements are judged by the slices and failure costs that matter, not only by an aggregate benchmark
- The company can distinguish a true gain from label noise, sample bias, leakage, evaluator error, or a shifted workload
- Changes that improve one stage cannot hide escaped defects, over-redaction, utility loss, or review burden somewhere else
- At least one consequential decision changes because the evidence reveals a risk, tradeoff, or opportunity that was previously unclear
- Evaluation becomes faster and more repeatable without sacrificing independence or rigor
- Quality claims communicate uncertainty honestly and remain understandable to engineers, customers, and risk owners
You Might Thrive Here If
- You have at least three years of professional experience in applied science, data science, machine learning, quantitative research, or a closely related role
- You have designed evaluations or experiments that changed a product, model, release, or operational decision
- You understand sampling, uncertainty, precision, recall, F1, calibration, agreement, class imbalance, distribution shift, and imperfect labels
- You can investigate messy, multi-stage data systems and determine where an apparent gain or loss actually came from
- You are comfortable writing Python and SQL and building reproducible technical artifacts rather than handing requirements to an engineering team
- You can protect the independence of an evaluation while collaborating closely with the people whose work it evaluates
- You have startup experience and enjoy broad ownership, changing context, and building the measurement foundation while decisions are already moving quickly
- You use AI tools fluently but do not confuse an articulate model output with valid evidence
- You communicate uncertainty and difficult findings directly, without hiding behind false precision
This Role May Not Be for You If
- You want to optimize models as your primary job rather than determine whether changes actually improve delivered data
- You prefer descriptive dashboards that stop short of changing a decision
- You treat labels, benchmarks, or model-based judges as ground truth without investigating how they fail
- You need a perfectly defined dataset and research plan before you can make progress
- You are uncomfortable disagreeing with a technically strong team when the evidence does not support its conclusion
- You do not want AI tools to be part of your daily scientific and technical workflow
Bonus
- Experience evaluating NER, entity resolution, information extraction, document understanding, multimodal, retrieval, or LLM systems
- Experience with privacy, de-identification, data quality, model risk, safety, or other high-trust decision systems
- Experience designing human-review, adjudication, weak-supervision, or active-learning systems
- Experience building adversarial corpora, replay systems, simulation environments, programmatic verifiers, or model-judge evaluations
- Experience connecting offline measures to escaped defects, customer outcomes, review effort, or preserved data utility
- Experience measuring quality across multi-stage batch or data pipelines
- ...transform their businesses and stay ahead of the curve in a rapidly evolving landscape. The Role: THE COMPANY is seeking a Lead Data Scientist to join our team and drive advanced data science solutions for one of our clients. This is a client-facing leadership role,...SuggestedContract workRemote work
$140k - $190k
...difference our hard work makes, and continue on our own paths of lifelong learning. How can you make an impact? As a Lead Data Scientist, you will join a team of data scientists, AI researchers, psychometricians, and software developers who provide technical and...SuggestedRemote workWorldwide$150k - $195k
...Lead Data Scientist Prism Data is building the future of credit risk assessment using modern data science and transaction-level financial data. Our API-based platform enables banks, fintechs, and lenders to use automated cash flow underwriting—analyzing detailed banking...SuggestedWork experience placementLocal area- ...paycheck to paycheck access to more affordable financial services and getting them on the path to better financial wellness. As a Data Scientist on our team you'll be responsible for building, improving and maintaining the key ML models that enable our services. The key...SuggestedLocal areaFlexible hoursShift work
- ...clients.Currently, we are looking for entry-level software programmers, Java full stack developers, Python/Java developers, data analysts/data scientists, machine learning engineers for full time positions with clients. Who should apply? Recent computer science/engineering/...SuggestedFull timeH1bImmediate startRemote work
$161.8k - $184.6k
...Principal Data Scientist - AI Foundations, Specialist Models Data is at the center of everything we do. As a startup, we disrupted the credit card industry by individually personalizing every credit card offer using statistical modeling and the relational database,...Full timePart timeLocal areaImmediate startFlexible hours- ...education in underserved communities and helping organizations achieve their full potential with AI. About the Role A Lead Data Scientist is responsible for designing and implementing data-driven solutions to complex business problems. The role requires extensive...Local area
- ...Job Title: AI/ML Data Engineer (Python, GenAI, ML Modeling) Location: New York, NY (Hybrid - 3 Days Onsite) Duration:... ...decision-making. This role requires collaboration with data scientists, engineers, product teams, and business stakeholders to design...Contract work
- About the Role We’re seeking a future team member for the role of Director, AI / Machine Learning Engineer to join our AI Hub team. This role is located in New York, NY. This is a highly visible leadership opportunity for a hands‑on technical leader who can combine...Work experience placement
$186k - $222k
...the team: architecture decisions, code quality, testing practices, and deployment patterns. You will work across the full stack—from data pipelines and ML models to LLM orchestration and Salesforce integrations—with the options to choose the right tool for each problem...Temporary work$110k - $169k
...for real work: We offer generous compensation packages that recognize hard work and excellence. Job Description The Pricing Data Engineer builds and maintains the data infrastructure and tools that enable consistent, data-driven pricing across the firm while modernizing...Full timeWork at office- ...responsible AI adoption. To be successful in this role, we're seeking the following: ~ Bachelor's degree in Computer Science, Data Engineering, or a related field or the equivalent combination of education and experience required. ~5-9 years of experience in...
- ...have you join us. The Role: Bedrock Robotics is hiring a Data Scientistto lead high-impact data science work across autonomy,... ...Bedrock. This role is ideal for a senior-to-staff-level data scientist who enjoys operating in ambiguous problem spaces, working close...
- ...Location (mandatory): New York, NY (Onsite) Apply knowledge of statistics, machine learning, programming, data modeling, simulation, and advanced mathematics to recognize patterns, identify opportunities, and make valuable discoveries leading to prototype biosensor development...Flexible hours
- ...Summary: This role involves building and delivering advanced data science and AI/ML solutions in an agile environment, with a focus... ...GenAI to accelerate outcomes. Supporting roles include Data Scientists and Data Engineers who will collaborate to build models, manage...
$100k - $130k
...Candid is a nonprofit that provides the most comprehensive data and insights about the social sector. We get you the information... ...need to do good. Candid currently has an opportunity for a Data Scientist. Candid (candid.org), the nation's leading authority on philanthropy...Temporary workSummer workLocal areaRemote workMonday to FridayFlexible hoursShift work$136k - $184k
...You'll partner closely with product, engineering, finance, and data teams to measure the impact of new features, design and analyze... ...decision-making You Have: ~3+ years of experience as a data scientist, applied scientist, economist, or related field; OR a PhD in...Flexible hours$200k - $225k
...talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire. Data Scientist Location - New York, NY / Williamsburg, Brooklyn On-site role requiring five days per week in-office in Williamsburg,...Work at officeRemote workVisa sponsorship$157.3k - $212.8k
...You'll partner closely with product, engineering, finance, and data teams to measure the impact of new features, design and analyze... ...Employee Discount Basic Qualifications ~3+ years of data scientist experience Preferred Qualifications Master's or PhD in...Local areaFlexible hours$125k - $187k
...millions of job seekers experience. Responsibilities Oversee data-driven initiatives that uncover opportunities to improve... ...and long-term product strategies. Mentor and guide other data scientists, fostering best practices in methodology, modeling, and analytical...Temporary workWork experience placementLocal area- ...Data Scientist As a Data Scientist on our team, you will not only define and predict our company's core metrics, you will drive them. This is a role of immense range, stretching from developing advanced analytical models to owning end-to-end strategic project implementation...Full timeWork experience placementWork at office
- ...Data Scientist The NYC Department of Consumer and Worker Protection (DCWP) protects and enhances the daily economic lives of New Yorkers to create thriving communities. DCWP licenses nearly 45,000 businesses in more than 40 industries and enforces key consumer protection...
- ...Data Scientist Job Title Data Scientist Location [Location / Remote / Hybrid] Employment Type Full-time Job Summary We are seeking a motivated Data Scientist to analyze complex datasets, build predictive...Full timeRemote workFlexible hours
- ...be modeling price elasticity to determine how a 5% discount impacts contribution margin. The next, you'll be digging into referral data to engineer a viral loop that lowers our CAC by half. You will hunt down the "why" behind every chart, discovering the invisible friction...Work at officeLocal area
$72.39k - $109.48k
...undergone a profound transformation by scaling a new model connecting data, creativity, and technology. The Groupe has continued this... ...center of the Groupe. Made up of audience strategists, data scientists, and analytics practitioners, Groupe Solutions will partner with...Temporary workFreelanceFlexible hoursShift work$100k - $135k
...Position: Data Scientist Who we are: Mediacom Communications Corporation is the 5th largest cable operator in the United States and the leading gigabit broadband provider to smaller markets primarily in the Midwest and Southeast. With a team of over 3,9...Work experience placementLocal areaFlexible hours- .... You're not here to run reports or maintain someone else's dashboards. You're here to turn ambiguous product questions into clear, data-backed answers that directly shape what we build and how we grow. You'll work alongside engineers and product teams, using AI to move...Work at officeRelocation packageFlexible hours3 days per week
- ...Applied Physics is seeking a Data Scientist experienced with a diverse array of data types to join our dynamic and multidisciplinary team of independent and entrepreneurial computer scientists and engineers. In this role, you will collaborate with scientists and researchers...Flexible hours
- ...Data Scientist The Data Scientist designs, develops, and implements advanced analytics, artificial intelligence (AI), and machine learning solutions that improve business decision-making and operational performance. The Data Scientist partners with business stakeholders...
- ...About the job Data Scientist Data Science is at the core of Online's business. Our team of researchers come from diverse disciplines and they drive innovation, new product ideation, experimental design and testing, complex analysis and delivery of data insights...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Scientist. Be the first to apply!
- entry level data scientist remote New York, NY
- data scientist New York, NY
- associate data scientist New York, NY
- principal data scientist New York, NY
- python data scientist (contract) New York, NY
- part time data scientist New York, NY
- data scientist machine learning engineer New York, NY
- data scientist (hedge fund) New York, NY
- ai data scientist New York, NY
- work from home data scientist New York, NY


