engineering
Posted 16 hours agoLead Engineer, Data Engineering
India Bangalore, IndiaOn-site
Responsibilities
- Implement scalable batch and near real-time data pipelines using PySpark, Python, and SQL on cloud-native platforms.
- Translate architecture and design patterns from the Data Engineering Principal into concrete, well-structured workflows and reusable components.
- Develop and maintain Airflow DAGs, managing dependencies, schedules, and operational monitoring for robust, fault-tolerant pipelines.
- Build data quality checks into pipelines (schema validation, record-level rules, thresholds, anomaly detection), and surface DQ metrics via dashboards or alerts.
- Implement data transformations and business rules to support use cases such as targeting, incentive compensation, call planning, patient journey, and market access analytics.
- Collaborate with Analytics, Data Science, and Product teams to understand
Requirements
- Work across data platforms such as Databricks and Snowflake, optimizing jobs through partitioning, clustering, caching, and cost/performance tuning.
- Strong, hands-on PySpark skills for ETL/ELT, including working with large datasets, optimization (partitioning, joins, caching), and troubleshooting performance issues. Solid Airflow
- experience with core life sciences commercial datasets: claims (medical/pharmacy), prescription (TRx/Nrx), sales, affiliations/rosters, call activity, and payer/plan/formulary data.
- Strong SQL skills (complex joins, window functions, CTEs, performance tuning) across cloud data warehouses such as Snowflake, Redshift, BigQuery, or Databricks SQL.
- Experience implementing data quality checks and validation patterns (row counts, referential integrity, business rules, reconciliation against source systems).
- Familiarity with modern data platforms (Databricks, Snowflake, Delta Lake, S3/ADLS/GCS) and data lakehouse concepts.
- Experience with at least one major cloud provider (AWS, Azure, or GCP) and core services used in data pipelines (e.g., S3/ADLS, Glue/Data Factory, Lambda/Functions).
- Comfort working in Git-based workflows and using CI/CD for data pipelines.
- Solid understanding of dimensional modeling and how to design fact and dimension tables for reporting and analytics.
- experience with pharma/life sciences commercial analytics (field force effectiveness, patient journey, HUB/SP, market access, or payer analytics).
- Experience with dbt or similar tools for modular SQL transformations and documentation.
- Exposure to data observability tools (e.g., Great Expectations, Soda, Monte Carlo) and data lineage/metadata tools.
- Familiarity with Kafka or Kinesis for streaming use cases.
- Experience containerizing data workloads (Docker) and running them on Kubernetes or similar orchestration platforms. Prior consulting or client-facing work where you gathered
Experience
- experience building production-grade data engineering solutions, with at least 3+ years hands-on with PySpark and Airflow.
Benefits
- Clear, concise communication and the ability to work with cross-functional stakeholders (consultants, analysts, data scientists) in a fast-paced environment. Bonus Points Deeper
Additional details
- Ingest, normalize, and harmonize life sciences commercial datasets (e.g., claims, prescription, sales, roster/territory, CRM, specialty pharmacy, payer/plan, formulary) into curated layers.
- Contribute to and follow data modeling standards (dimensional models, star/snowflake schemas) aligned to life sciences commercial subject areas.
- Embed testing (unit tests for transformations, integration tests for pipelines, and regression checks on key metrics) into the development lifecycle.
- requirements and ensure datasets are usable, documented, and trusted Produce clear technical documentation for pipelines, schemas, and business logic, making it easy for others to extend and maintain your work.
- Participate in Agile ceremonies, provide realistic estimates, and own stories from design through to production support and handover.
- experience: DAG design, scheduling, sensors, operators, and managing operational run health. Practical
- Strong debugging skills, willingness to dive into logs and metrics, and bias toward shipping working solutions quickly and iterating.
- requirements and translated them into concrete data solutions.