About the Position
Introduction
Natera is seeking an experienced Senior Software Engineer with a deep scientific R&D background in modern data engineering and AI-enabled development. This role focuses on designing and building data products that directly support genomics research and translational science. The ideal candidate will possess strong data engineering skills, a computer science background, and hands-on experience in bioinformatics, genomics, or computational biology, with the ability to work independently in an R&D environment. You will prototype novel data products and ensure they evolve into robust, compliant, and scalable platforms, emphasizing reproducibility, traceability, performance, and scientific usability.
Responsibilities
- Design, build, and maintain data products supporting R&D, analytics, Lab, and scientific workflows, from initial design through deployment and iterations.
- Build and maintain data pipelines for large and complex datasets, from raw inputs through derived and analysis-ready datasets.
- Apply domain knowledge in genetics and bioinformatics to design data models, schemas, and abstractions aligned with research patterns and downstream analysis.
- Design and enforce de-identification and privacy-preserving architectures that meet HIPAA and related regulatory requirements while remaining usable for research.
- Design scalable data models for analytics, reporting, and downstream applications, maintaining high standards of data quality, accuracy, lineage, and observability.
- Partner closely with R&D scientists, bioinformatics teams, and software engineers to translate research needs into well-structured, reusable data assets.
- Optimize storage, retrieval, and lifecycle management for large scientific files (e.g., sequencing data, intermediate artifacts, derived datasets).
- Drive rapid prototyping efforts for exploratory, proof-of-concepts, and early-stage initiatives, guiding the transition to production-grade systems.
- Implement best practices for data quality, validation, lineage, observability, and reproducibility.
- Collaborate with product managers and domain experts to translate requirements into technical solutions.
- Establish golden paths (templates, examples, docs) and contribute to shared data product catalogs, patterns, and best practices.
- Provide technical guidance and mentorship to mid-level engineers.
Requirements
- Bachelor’s or Master’s degree in computer science or bioinformatics, with healthcare or biotech data domain experience preferred.
- 8+ years of experience in data engineering, designing and maintaining data pipelines and cloud data architectures (e.g., Snowflake, AWS, etc.).
- Strong background in bioinformatics, genomics, or computational biology (required), understanding key genomics and bioinformatics data formats (BAM, VCF, FASTQ), common compression techniques, and their storage, delivery, and management needs.
- Demonstrated experience supporting scientific R&D, Lab workflows, and research teams with production-grade data systems.
- Strong proficiency in Python, SQL, and distributed processing frameworks (Spark or equivalent).
- Experience with modern orchestration tools (Airflow, dbt, Dagster).
- Experience leveraging AI-assisted development tools (e.g., LLM copilots) to accelerate data solution development.
- Familiarity with building data products that support analytics, ML, or AI applications.
- Strong data modeling expertise (dimensional, normalized, healthcare-specific schemas).
- Experience implementing CI/CD for data pipelines and IaC (Terraform, CloudFormation); Knowledge of data observability, testing, and data quality frameworks.
- Demonstrated ownership of production-grade data systems and end-to-end pipeline lifecycle.
- Ability to evaluate emerging data and AI technologies and recommend scalable solutions.
- Proven ability to operate effectively in fast-paced environments, balancing speed, rigor, and compliance.
- Strong written and verbal communication skills with the ability to collaborate across engineering, analytics, and business stakeholders.
- Experience working with healthcare, life sciences, or other highly regulated data, including hands-on HIPAA compliance.
Nice to Have
- Exposure to vector databases, embeddings, semantic search, or RAG-based architectures.
Benefits
[Benefits not specified in the provided job description]
About Company
[About Company information not specified in the provided job description]
How to Apply
Please apply through the provided application link.
Apply Now
Your data is only shared with Natera
Location
US Remote
Type
Full-Time
Keywords
Similar Roles
Explore comparable positions
Senior Software Engineer, Data & AI Solutions
Natera