Job Description
Job DescriptionDescription:
Job Summary
The Senior Data Engineer – Data Operations designs, builds, and maintains the pipelines and infrastructure that move healthcare data (claims, clinical, eligibility, financial, census, biometrics) reliably across systems. This role owns the technical backbone of data operations — ensuring pipelines run on schedule, data is validated and trustworthy, and downstream analytics, sales, and operations teams have consistent access to accurate data. The ideal candidate combines strong engineering fundamentals with healthcare domain awareness and an operations mindset focused on reliability and scale.
Responsibilities:
• Design, build, and maintain scalable ETL/ELT pipelines to ingest and process healthcare data (claims, EHR/EMR, eligibility, enrollment, Census, biometrics) from multiple source systems
• Own the technical health of production data pipelines — monitoring, alerting, incident response, and root-cause resolution for failures or data delays
• Architect and maintain data models, warehouses/data lake and schemas that support operational reporting, analytics, and compliance reporting needs
• Implement automated data quality checks, validation rules, and reconciliation processes to catch errors before they reach downstream consumers
• Optimize pipeline performance, storage costs, and query efficiency across cloud data platforms
• Partner with business/data analysts, operations, and compliance teams to translate operational and reporting requirements into robust data engineering solutions
• Ensure all data handling complies with HIPAA, HITECH, and relevant payer/state/federal data security and privacy regulations
• Build and maintain CI/CD pipelines for data infrastructure, versioning schema changes, and automating deployments
• Document data lineage, pipeline architecture, and system dependencies to support audits, onboarding, and operational continuity
• Support integration of new data sources (payer feeds, clearinghouses, third-party vendors) into existing data operations infrastructure
• Contribute to engineering best practices, code review standards, and technical documentation
• Evaluate and recommend new tools, frameworks, or AI-assisted engineering practices to improve pipeline reliability and team productivity
Requirements:
Required Qualifications:
• Bachelor's degree in Computer Science, Data Engineering, Information Systems, or related field
• 5+ years of data engineering experience, with at least 3 years in healthcare, health insurance, or managed care data environments
• Strong proficiency in SQL and at least one programming language (Python, Scala, or Java) for pipeline development
• Hands-on experience building and maintaining ETL/ELT pipelines using tools such as Airflow, dbt, Informatica, or Glue
• Experience with cloud data platforms (Snowflake, AWS Redshift/S3, Azure Synapse/Data Lake, GCP BigQuery, or Databricks)
• Working knowledge of healthcare data standards and formats: HL7, X12 EDI (837/835/834), FHIR, ICD-10, CPT, HCPCS, DRG
• Solid understanding of data warehousing concepts, dimensional modeling, and data lake/lakehouse architectures
• Familiarity with HIPAA and healthcare data privacy/security compliance requirements
• Experience with version control (Git) and CI/CD practices for data infrastructure
• Strong troubleshooting skills with the ability to diagnose and resolve complex pipeline and data quality issues under time pressure
Preferred Qualifications:
• Experience with orchestration and monitoring tools (Airflow, Prefect, Dagster) and observability platforms (Monte Carlo, Datadog)
• Familiarity with infrastructure-as-code (Terraform, CloudFormation) and containerization (Docker, Kubernetes)
• Experience with streaming data technologies (Kafka, Kinesis, Spark Streaming) for near-real-time healthcare data feeds
• Exposure to master data management (MDM) and data governance frameworks
• Hands-on experience using AI coding assistants (e.g., Claude, GitHub Copilot) to accelerate pipeline development, code review, and documentation
• Familiarity with AI Development Life Cycle (AIDLC) practices for incorporating AI-assisted workflows into engineering and data operations processes
• Certification in a major cloud platform (AWS, Azure, or GCP) or relevant data engineering credential
Key Competencies:
• Strong engineering fundamentals and systems thinking
• Operational reliability mindset — building for scale, monitoring, and failure recovery
• Attention to detail in data quality and compliance-sensitive environments
• Cross-functional collaboration with analytics, operations, and compliance teams
• Ownership and accountability for production systems
• Clear technical communication and documentation habits
