Infosys
Infosys

PySpark Developer

직무데이터 엔지니어링
경력중급
위치Bangalore, India
근무오피스 출근
고용Technology Analyst
게시2개월 전
지원하기

포지션 소개

Key Responsibilities:

Develop and maintain data pipelines using Py Spark
Process and analyze large-scale datasets in distributed environments
Design and implement ETL/ELT workflows
Optimize Spark jobs for performance and scalability
Work with data stored in HDFS, Hive, or cloud storage (S3, ADLS)
Collaborate with data engineers, analysts, and business teams
Ensure data quality, integrity, and governance
Debug and troubleshoot data processing issues
Automate workflows using scheduling tools (Airflow, Oozie, etc.)
Write clean, scalable, and efficient code

Required Skills & Qualifications:

Technical Skills:

Strong proficiency in Python and Py Spark:

Good experience with Apache Spark (RDDs, Data Frames, Spark SQL)
Knowledge of Hadoop ecosystem (HDFS, Hive)
Experience in ETL pipeline development
Familiarity with SQL and database concepts
Experience with data formats (Parquet, ORC, JSON, CSV)
Basic understanding of distributed computing concepts
Exposure to version control tools (Git)

  • Technology->Big Data
  • Data Processing->Py Spark

Education: MCA,MSc,MTech,Bachelor of Engineering,BCA,BSc,BTech

  • Preferred skills: Technology->Big Data
  • Data Processing->Py Spark

Infosys 소개

BANGALORE

본사 위치