Search by job, company or skills

Data Engineer - Gen AI & LLM

Early Applicant
  • Posted 20 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

We are looking for a passionate Data Engineer with 5+ years of experience in designing,

developing, and maintaining scalable data platforms and ETL/ELT pipelines. The ideal candidate

should possess strong expertise in Python, SQL, cloud data services, and modern data

engineering frameworks. You will play a key role in building reliable, high-performance data

solutions that support analytics, reporting, and AI/ML initiatives.

Key Responsibilities

• Design, develop, and maintain scalable data pipelines and ETL/ELT workflows.

• Build and optimize data ingestion processes from multiple structured and unstructured

data sources.

• Develop robust data models and data warehouses for analytics and reporting.

• Design and optimize SQL queries for high-performance data processing.

• Build and maintain data lakes using cloud storage solutions.

• Implement data validation, cleansing, transformation, and quality checks.

• Integrate data solutions with cloud platforms such as AWS, Azure, or Google Cloud

Platform.

• Develop batch and real-time data processing pipelines using modern data processing

frameworks.

• Collaborate with data scientists, analysts, and application teams to deliver reliable data

solutions.

• Follow software engineering best practices, including Git, CI/CD pipelines, testing,

monitoring, and technical documentation.

Required Skills

• 5+ years of experience in Data Engineering.

• Strong programming experience in Python.

• Hands-on experience with PySpark and Apache Spark.

• Advanced SQL skills, including query optimization and performance tuning.

• Experience with AWS services/ Azure/ GCP.

• Experience building scalable ETL/ELT pipelines.

• Strong exp in Generative AI and Large Language Models (LLMs)

• Knowledge of data warehousing concepts and dimensional modeling.

• Experience working with large-scale distributed datasets.

• Familiarity with Git and CI/CD pipelines.

• Understanding Agile/Scrum methodologies.

Preferred Qualifications

• Experience with Apache Airflow or similar workflow orchestration tools.

• Knowledge of Kafka and real-time data streaming.

• Familiarity with Docker, Kubernetes, and Terraform (IaC).

• Exposure to Databricks and modern data lake technologies (Delta Lake, Apache Iceberg,

or Apache Hudi).

• Understanding data governance, metadata management, and CI/CD practices.

• Strong exp in Generative AI and Large Language Models (LLMs

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153699269

Beware of Scammers

We don’t charge money for job offers