Machine Learning Engineer - Clinical NLP Job at Novellia, Remote

OU41NmhsSDRUVXNQcjY3bm1jSjZmdm54WkE9PQ==
  • Novellia
  • Remote

Job Description

About the role

Most of what matters in a health record isn't in a structured field - it's in the note, the discharge summary, the pathology report, the scanned fax. Turning that unstructured clinical text into trustworthy, structured features is what makes a longitudinal health history usable for research, and it's one of the highest-leverage capabilities Novellia can own.

You'll be our first ML hire, joining Platform Engineering and reporting to the Head of Platform Engineering, as technical owner of this multi-quarter effort. The interesting decisions are still open - what we extract first, how we know we're right, what a mature extraction pipeline looks like at our scale. There's no existing approach to inherit or defend.

The work draws on two toolkits. Roughly 70% is applied ML on clinical text: entity extraction, classification, sequence labelling, annotation strategy, error analysis, calibration, and the evaluation discipline that tells you whether your numbers mean anything. Roughly 30% is LLM-based: prompt development, structured output, retrieval, and the evals and observability that keep generative approaches honest. Deciding which approach a given problem calls for is the most interesting part of the job, and that call is yours.

We're looking for a leader in this seat: setting technical direction rather than waiting to be handed a problem. If this grows the way we think it will, leading the team we build around it is on the table.

What you'll do

  • Own the full lifecycle of extraction models - framing, data/annotation strategy, model selection, training/fine-tuning, evaluation, deployment, monitoring, retraining. Not a research seat, not a hand-off seat.
  • Define what "accurate enough" means with clinical and customer-facing stakeholders, and build the evaluation harness that makes the answer defensible - the first deliverable, not a follow-up.
  • Partner with Clinical Data Managers on curation design and own the technical half of QA/QC alongside them: which variables are extractable, how an instruction becomes a model spec, and the tooling/sampling/error analysis behind human-in-the-loop review.
  • Build clinical NLP pipelines against messy real-world data and work with backend engineers to productionize what you build.
  • Use LLMs with the same rigor you'd apply anywhere: versioned prompts, real evals, tracked cost/latency, known failure modes.
  • Make extraction quality legible to non-ML colleagues, and treat de-identification, PHI handling, audit trails, and access controls as part of the modelling problem, not someone else's checklist.
  • Help shape the roadmap around the problems you see - a mission and a close working partner, not a backlog.

What we're looking for

  • Healthcare or life sciences experience with real clinical data - clinical notes, EHR data, claims, registries, or similar. This one is not negotiable for us.
  • 6+ years in applied ML , with models you personally took from problem statement to production and kept working - you know what degraded, how you found out, and what you did.
  • Depth in applied ML on text: information extraction, NER, classification, sequence labelling, weak supervision, and the evaluation practice around them, including annotation guidelines and inter-annotator agreement you've had to act on.
  • Practical, current experience with LLM-based approaches: prompt development, structured output, retrieval, fine-tuning where warranted, evals and observability for generative systems - enough to know where they help, and where they quietly don't.
  • Strong engineering fundamentals in Python . Your work runs in production, not only in a notebook.
  • Strong collaboration instincts across the ML boundary: you define problems with stakeholders before solving them, write clearly, and bring people along.
  • A track record of solving problems rather than closing tickets. Self-directed, comfortable without a playbook, and comfortable being wrong in public when the evidence says so.

Nice to have

  • Fluency with clinical terminologies and standards: SNOMED CT, ICD-10, LOINC, RxNorm, CPT, FHIR
  • Experience with HIPAA, SOC 2, de-identification methodology, or IRB and regulatory-grade data work
  • Experience as an early or first ML hire
  • Experience building or running human-in-the-loop annotation and QC operations at scale
  • OCR and document-understanding experience on low-quality real-world documents
  • Experience mentoring or leading ML engineers, or interest in growing that way

What this role is not

  • Not a research role. The bar is extraction quality in production, not publications.
  • Not an LLM-wrapper role. If your instinct is that every problem is a prompt away from being solved, we'll frustrate each other.
  • Not a large-team role yet. You'd be the first ML engineer in a small Platform Engineering function - breadth and influence, and fewer specialists to lean on.
  • Not a role where someone hands you a clean labelled dataset. Building it is the job.

Why this role is a good bet

  • Ground-floor ownership of a capability with direct commercial weight, with influence over architecture, roadmap, and eventually hiring.
  • Both halves of the modern ML toolkit in one seat, on a problem where the choice between them genuinely matters.
  • A manager who treats process and people work as legitimate engineering work, and intends for this seat to grow.
  • Health tech means the work has stakes - better extraction means higher quality research

Benefits & Perks

  • Equity in Novellia
  • Medical, dental, and vision coverage
  • 401(k)
  • Flexible time off
  • Wellness stipend
  • Up to 12 weeks of parental leave

Don't meet every requirement? Studies show women and people of color are less likely to apply unless they meet every qualification. If you're excited about this role but your experience doesn't align perfectly, we encourage you to apply anyway - you may be the right fit for this or another role.

Job Tags

Full time, Flexible hours

Similar Jobs

EssilorLuxottica Group

Target Optical - Assistant Manager Job at EssilorLuxottica Group

 ...Requisition ID:934543 Store # :002101 Target Optical Position: Full-Time Total Rewards: Benefits/Incentive Information At Target Optical, we love the neighborhoods we belong to and thats why we care for them. By listening and building relationships... 

ReqRoute,Inc

Mainframe Developer Job at ReqRoute,Inc

 ...Mainframe Developer Location: Chicago, IL - 100% Onsite Duration: 6 months (contract) Experience Required: 8 10+ years Position Overview We are seeking an experienced Mainframe Developer to design, develop, maintain, and enhance applications... 

C-Fam

Assistant Director of Legal Studies Job at C-Fam

 ...research and analysis of UN documents, state and federal law, international law,and foreign law. Identify opportunities to engage in UN...  ...written interventions Draft statements, reports, memos, white papers, legal briefs, and other submissions forinternational bodies... 

Confidential

PE/Sports Teacher Football/Basketball Job at Confidential

Position TitlePE/Sports Teacher Football/Basketball/Volleyball/Swimming/Rugby The Role Deliver physical education programmes specialising...  ...What Were Looking For Bachelors degree in PE, Sports Science, or related field PE certificate for z visa application ... 

Shield Consulting Solutions

Backend Application Engineer (Python) Job at Shield Consulting Solutions

 ...and Responsibilities: The Application Engineer (Backend) will provide the development, test, deploy, and sustainment of various Python based ReST end points, microservices, and data model management capabilities utilizing Django and Flask frameworks to interact with...