Infosys
Infosys

Data Engineer

职能数据工程
级别中级
地点Pune, India
方式现场办公
类型Senior Consultant
发布2个月前
立即申请

职位介绍

Role demands a highly skilled Data Engineer to design, build, and optimize scalable data pipelines and data platforms. The ideal candidate will have strong expertise in data modeling, cloud-based data architectures, and modern data engineering tools across Azure, Snowflake, and Databricks environments.

Key Responsibilities:

  • Data Engineering & Pipeline Development

  • Design, develop, and maintain robust ETL/ELT pipelines using Databricks, PySpark, and Azure Data Factory (ADF).

  • Build scalable and efficient data ingestion frameworks for structured and unstructured data.

  • Optimize pipeline performance through performance tuning and orchestration best practices.

  • Data Modeling & Management

  • Develop and maintain data models using modern tools (DBT preferred).

  • Implement Master Data Management (MDM) solutions to ensure data consistency and integrity.

  • Design scalable and efficient Snowflake schemas (star/snowflake schema, dimensional modeling).

  • Database & Query Optimization

  • Write and optimize advanced SQL queries across Snowflake, Azure SQL, and Synapse.

  • Develop and manage stored procedures and database objects.

  • Ensure efficient data retrieval through indexing, partitioning, and query optimization.

  • Cloud & Platform Integration

  • Work with Azure data services including:

  • Azure Data Factory (ADF)

    • Azure Data Lake Storage (ADLS)
    • Azure Synapse Analytics
    • Azure SQL Database
  • Integrate and maintain Snowflake with Azure ecosystem.

  • Python Development

  • Develop data transformation and automation scripts using Python libraries:

  • pandas

    • pyodbc
    • SQLAlchemy
  • Build reusable components for data processing and validation.

  • Data Quality, Validation & Monitoring

  • Implement data validation rules, quality checks, and anomaly detection frameworks.

  • Perform root cause analysis for data inconsistencies.

  • Develop dashboards or tools for data quality monitoring.

  • Collaboration & DevOps

  • Use GitHub for version control, branching strategies, and code reviews.

  • Manage workload scheduling and dependency management for pipelines.

  • Collaborate with cross-functional teams including data analysts, data scientists, and business stakeholders.

  • Required Skills & Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Information Systems, or related field.

  • Strong experience in data engineering and data platform development.

  • Technical Skills

  • Expertise in DBT (preferred) for data modeling.

  • Strong SQL skills with hands-on experience in:

  • Snowflake

    • Azure SQL
    • Stored procedures
  • Proficiency in Python for data engineering workflows.

  • Hands-on experience with:

  • Databricks & PySpark

    • Azure Data Services (ADF, ADLS, Synapse)
  • Strong knowledge of Snowflake architecture and schema design.

  • Experience with data validation, quality frameworks, and analysis tools.

  • Familiarity with GitHub and CI/CD practices.

  • Preferred Qualifications

  • Experience implementing data governance and MDM solutions.

  • Knowledge of performance tuning in distributed processing systems.

  • Familiarity with workflow orchestration tools.

  • Experience in agile environments (SCRUM/Kanban).

Education: Master Of Engineering,Master Of Technology,Bachelor of Engineering,Bachelor Of Technology

  • Preferred skills: Technology->Data Engineering->Databricks,Technology->Cloud Integration->Azure Data Factory (ADF),Technology->Big Data
  • Data Processing->Py Spark,Technology->Oracle->PL/SQL,Technology->Data on Cloud-Data Store->Snowflake,Technology->Cloud Platform->Databases on Azure->Azure SQL Database,Technology->Cloud Platform->Azure Analytics Services->Azure Data Lake

关于Infosys

PUNE

总部位置