Job Overview
Mphasis is looking for a Senior Big Data / ETL Engineer with 5–8 years of experience in designing, developing, and supporting enterprise-scale Big Data, ETL, and streaming solutions.
The role requires strong hands-on experience with Hadoop, PySpark, Python, HDFS, Kafka, Streaming, ETL, and Informatica, along with exposure to Azure and/or AWS, DevOps, CI/CD, and AI-assisted development tools such as Microsoft Copilot.
Job Details
| Category | Details |
|---|---|
| Company | Mphasis |
| Job Role | Senior Big Data / ETL Engineer |
| Experience | 5–8 Years |
| Location | Bangalore |
| Work Model | Hybrid – Minimum 3 days/week from RMZ Ecoworld, Bangalore |
| Cloud | Azure and/or AWS |
| Employment Type | Full-Time |
| Primary Skills | Hadoop, PySpark, Python, HDFS, Kafka, Streaming, ETL, Informatica, DevOps, Copilot, Markdown |
Key Responsibilities
Big Data Engineering
- Design, develop, test, and support scalable data processing solutions using Hadoop ecosystem technologies.
- Build and optimize batch and real-time data pipelines using PySpark and Python.
- Work extensively with HDFS, distributed computing, and large-volume data processing.
- Tune Spark and Hadoop workloads for performance, scalability, and reliability.
- Implement monitoring, error handling, reconciliation, and recovery mechanisms.
ETL & Informatica
- Design and support ETL workflows for enterprise data ingestion, transformation, and integration.
- Work with Informatica mappings, workflows, sessions, and job monitoring.
- Perform data validation, reconciliation, dependency tracking, and failure analysis.
- Collaborate with source-system, data warehouse, and reporting teams.
Kafka & Streaming
- Design and implement event-driven data pipelines using Apache Kafka and streaming frameworks.
- Develop producer and consumer applications.
- Work with Kafka topics, partitions, replication, and retention.
- Troubleshoot streaming workloads and maintain reliable real-time data flows.
- Implement validation and data quality checks for streaming pipelines.
Cloud Engineering
- Develop and deploy Big Data, ETL, and streaming solutions on Azure and/or AWS.
- Work with cloud storage, compute, monitoring, and security capabilities.
- Support hybrid and cloud-based data architectures.
DevOps & Automation
- Build and maintain CI/CD pipelines for data engineering and ETL applications.
- Use Git-based version control, branching, and release management.
- Automate build, testing, deployment, and operational workflows.
- Apply Docker and Kubernetes fundamentals where applicable.
- Participate in production support, incident response, and root-cause analysis.
AI-Assisted Development & Documentation
- Use Microsoft Copilot or equivalent AI tools for code generation, optimization, debugging, and unit-test creation.
- Create and maintain Markdown (.md) documentation.
- Prepare technical designs, data-flow documentation, runbooks, SOPs, release notes, and troubleshooting guides.
- Maintain accurate documentation in version-controlled repositories.
Required Technical Skills
| Skill Area | Expected Skills |
| Big Data | Hadoop, HDFS, Spark, PySpark, Hive preferred, distributed computing |
| ETL & Informatica | ETL design, Informatica, mappings, workflows, sessions, monitoring, troubleshooting, reconciliation |
| Programming | Python, Shell/Bash scripting, SQL |
| Streaming & Messaging | Apache Kafka, Spark Structured Streaming, Kafka Streams, event-driven architecture |
| Cloud | Azure and/or AWS, cloud storage, data engineering services, monitoring and observability |
| DevOps | Git, CI/CD, Azure DevOps, GitHub Actions, Jenkins, Docker, Kubernetes fundamentals |
| Databases | Relational databases, SQL, NoSQL preferred |
| Documentation | Markdown, technical design documents, runbooks, SOPs, reusable engineering documentation |
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 5–8 years of relevant experience in Big Data, Data Engineering, ETL Engineering, or Streaming Engineering.
- Proven experience delivering enterprise-scale data pipelines and supporting production workloads.
- Hands-on experience with PySpark, Python, Hadoop/HDFS, Kafka, and ETL tools.
- Strong analytical, troubleshooting, and communication skills.
Ideal Candidate
The ideal candidate should have strong hands-on experience building and supporting large-scale Big Data and ETL pipelines, with practical knowledge of PySpark, Hadoop, Kafka, Informatica, cloud platforms, and DevOps practices. Experience with AI-assisted development and maintaining structured technical documentation will be an advantage.