Search by job, company or skills

Principal Data Architect

  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities

  • Ownership of Data Warehouse models and curation (Snowflake, Redshift, DBT)
  • Design, build, and optimize data pipelines using PySpark and Python for scalability and reliability
  • Implement and manage workflow orchestration with Airflow for scheduling and automation
  • Manage infrastructure-as-code with Terraform to ensure reproducibility and reliability of data environments
  • Build and manage serverless and distributed data processing workflows using AWS Glue, EMR, and Lambda
  • Drive the evolution of the data environment to deliver high-quality data, speed, and availability
  • Curate source-system data — including RDBMS sources — to deliver trusted, analytics-ready datasets
  • Provide input and involvement on data cataloging and data management efforts
  • Own production ETL/ELT performance tuning and environment-level resource consumption and management
  • Drive the migration of POC pipelines to production-ready processes
  • Lead technical decision-making on data architecture and mentor junior/mid-level engineers
  • Collaborate with analytics, product, and business teams to ensure data solutions align with organizational needs

Qualifications

  • 10+ years of experience in data engineering and data warehouse development
  • Strong SQL development skills with expertise in performance tuning on RDBMS (Postgres, MySQL, or similar)
  • Mandatory experience with: Terraform, Airflow, Snowflake, Redshift, DBT, Python, and PySpark
  • Hands-on experience with AWS Glue, EMR, and Lambda for building scalable, distributed data pipelines
  • Bachelor's or Master's degree in Computer Science, Information Technology, or related field
  • Proven experience designing and delivering enterprise-scale data warehouses and marts to support business analytics
  • Experience developing data curation and integration processes, metadata management, and data quality initiatives
  • Experience with streaming and real-time data processing (Spark, Databricks optional)
  • Familiarity with ETL tools (Fivetran good to have)
  • Experience working with broader AWS services — Step Functions, S3, CloudFormation
  • Knowledge of CI/CD workflows and source control (GitHub Actions, GitLab)
  • Demonstrated experience leading data engineering initiatives or mentoring teams
  • More Info

    Job Type:
    Industry:
    Employment Type:

    About Company

    Job ID: 153466791

    Similar Jobs

    Hyderabad, India

    Skills:

    Cloud Technologiesautomation solutionsData ExtractionDatabase KnowledgeEhrAdvanced SqlGovernanceOracle databasesconversion frameworkshealthcare data exchange standardshealthcare data conversion practicesEMR platformsvalidation methodologiesclinical workflowshealthcare data structuresdata migration toolshealthcare interoperability conceptshealthcare information systemsLoadingenterprise data migration practicesmigration testing methodologiesTransformationdata quality managementETL processesMappingReconciliationNormalization

    Hyderabad, India

    Skills:

    data engineering snowflake Informatica CloudMetadata ManagementData ModelingData IntegrationELTData ArchitectureData GovernanceAWSData QualityGcpData WarehousingAzureEtlAirflowAnalytics SolutionsCloud Data PlatformsSalesforceCortex AICoalesceGenerative AIData Lakehouse PatternsdbtMaster Data Management

    Beware of Scammers

    We don’t charge money for job offers