- Dawar Consulting, Inc.
- South San Francisco, CA
- Full-Time
- 38 days ago
- $80–$85 / hour
AI/ML Engineer - NLP Scientist.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
AI/ML Engineer - NLP Scientist: our view in 3 lines...
- The Role:This role is for a senior AI and NLP specialist working on claim verification systems for scientific, clinical, and regulatory content.
- The Person:The person will build claim verification and evidence attribution pipelines, develop retrieval and contradiction detection systems, create evaluation datasets, and support traceable human-in-the-loop review workflows.
- Requirements:The ideal candidate has strong Python production engineering, NLP / LLM / Generative AI development, RAG, hybrid search, vector search, lexical retrieval, Natural Language Inference, and expert-labeled datasets.
About the role
We are seeking a Senior AI/ML Engineer to build an evidence-grounded AI capability that verifies generated claims against approved scientific, clinical, regulatory, and reference materials before human review. The system will retrieve relevant evidence, decompose claims into verifiable assertions, evaluate evidence support, and provide traceable decisions with citations. The system must recognize unsupported or contradicted claims and abstain rather than guess.
-
Build production-grade Python/NLP pipelines for claim verification and evidence attribution.
-
Develop hybrid retrieval using lexical and vector search to identify relevant evidence.
-
Implement claim decomposition, natural language inference (NLI), entailment, and contradiction detection.
-
Evaluate whether generated claims are genuinely supported by cited evidence.
-
Design confidence thresholds, abstention logic, escalation rules, and human-in-the-loop workflows.
-
Build evaluation datasets with expert annotation guidelines and measure inter-annotator agreement.
-
Track false approvals, false rejections, abstentions, and other error categories.
-
Develop traceable systems that allow decisions to be reconstructed based on model version, evidence, citations, and reviewer actions.
-
Work with Medical, Legal, Regulatory, and scientific stakeholders to translate review requirements into technical solutions.
Required Skills
-
Strong Python production engineering.
-
NLP / LLM / Generative AI development.
-
RAG, hybrid search, vector search, and lexical retrieval.
-
Natural Language Inference (NLI), entailment, contradiction detection.
-
Claim decomposition and evidence attribution.
-
LLM/model APIs and production evaluation frameworks.
-
AI/ML evaluation, benchmarking, and error analysis.
-
Human-in-the-loop AI, confidence scoring and abstention.
-
Experience with scientific, technical, regulatory, legal, or other high-stakes content.
-
Experience creating expert-labeled datasets and annotation guidelines.
-
Strong understanding of traceability, citations, and reproducible AI decisions.
Preferred Skills
-
Knowledge graphs and relationships between claims, evidence, references, products, and indications.
-
Deterministic rules + ML/LLM decision systems.
-
Pharmaceutical, biotech, healthcare, regulatory, legal, financial compliance, or scientific
publishing experience.
-
Familiarity with clinical studies, statistics, scientific literature, and citation practices.
-
Experience with LangChain, LlamaIndex, Hugging Face, PyTorch, or similar NLP/
LLM frameworks.

