

Corning is hiring an AI Data Engineer to support its Advanced Analytics platforms and enterprise-scale data initiatives in Pune. This role focuses on building reliable big-data ingestion pipelines, working with structured, semi-structured, and unstructured data, and preparing high-quality datasets for analytics and data science teams. Candidates with knowledge of Python, Java or Scala, Apache Spark, AWS, Delta Lake, ETL/ELT pipelines, and data lakes will find this opportunity particularly relevant. The position offers exposure to production engineering practices, cloud technologies, data validation, CI/CD, and large-scale enterprise data environments.
This opportunity places a Data Engineer at the center of Corning's advanced analytics ecosystem, where reliable data movement, validation, and preparation are essential for turning operational information into usable datasets for engineering, analytics, and data science teams.
π Enterprise Data Engineering Exposure
Corning operates across multiple technology-driven industries, including life sciences, optical communications, consumer electronics, automotive, displays, and solar technologies. The Data Engineer will contribute to platforms that collect and prepare information from different operational and process data sources.
The central responsibility is not simply moving data from one location to another. The engineer is expected to help create reliable, maintainable, and instrumented ingestion pipelines that can support advanced analytics projects across the organization.
Production Big-Data Pipeline Development
Candidates joining this role will work with platform engineers, domain experts, application developers, controls engineers, and data scientists. This creates exposure to the complete journey of enterprise dataβfrom the original source system through ingestion and validation to datasets that can be used for analytics.
βοΈ Building Reliable Data Pipelines
A major part of this position involves designing, testing, deploying, and maintaining production data ingestion pipelines. These pipelines will bring data from multiple operational systems into on-premises and cloud-based data lakes.
The work includes following established engineering practices rather than treating data pipelines as one-time scripts. Version control, automated deployment practices, testing, and maintainability are important parts of the role.
Python Java Scala Apache Spark Git
The role specifically values programming ability in at least one modern JVM language such as Java, Scala, or Kotlin, along with Python. This combination is useful because enterprise data engineering environments often require both large-scale processing capabilities and flexible scripting for automation.
π Data Validation and Profiling
Moving data successfully does not automatically mean that the data is useful. Corning expects the Data Engineer to work on automated data validation and profiling capabilities that help ensure reliable delivery.
This means understanding whether incoming datasets meet expected quality and structural requirements before they are used by downstream teams.
The engineer will collaborate with data source teams to define ingestion requirements and validate whether the landed data is acceptable for business and technical use.
The objective is to deliver data that domain experts and analytics teams can trust without repeatedly performing the same manual validation work.
π§Ή Cleansing and Enrichment Work
After data is ingested, it may require additional preparation before becoming useful for analytics. The Data Engineer will work with domain experts and data scientists to understand requirements for data cleansing and enrichment.
This can involve transforming raw information into datasets that are more consistent, complete, and suitable for downstream analysis.
The role emphasizes validating the final datasets with the teams that will actually use them. That creates an important connection between data engineering work and practical business or analytical requirements.
A strong data pipeline is valuable when the resulting dataset is reliable and directly usable by the people who depend on it.
βοΈ AWS and Cloud Platform Exposure
The position provides entry-level exposure to several AWS services used in modern enterprise data environments.
Relevant technologies mentioned for this role include:
Amazon S3 for storage and data lake environments
Amazon EC2 for computing infrastructure
AWS DMS for data migration
Amazon RDS for relational database services
Amazon EMR for large-scale data processing
Amazon Redshift for data warehousing
AWS Lambda for event-driven computing
Amazon DynamoDB for NoSQL workloads
Amazon CloudWatch for monitoring
AWS CloudTrail for activity tracking
AWS Exposure Data Lake Environment
Understanding how these services fit into a broader data architecture can strengthen a candidate's ability to work in cloud-based engineering teams.
π₯ Spark and Modern Lakehouse Technologies
Apache Spark architecture is one of the important technical areas for this role. Candidates are expected to have familiarity with Spark and technologies commonly associated with modern data lake environments.
Traditional Data Storage Modern Lakehouse Architecture
The job description specifically mentions experience or familiarity with:
Apache Spark S3 Parquet Delta Lake
Understanding these technologies together is particularly useful. Spark supports distributed processing, Parquet is commonly used for efficient analytical storage, and Delta Lake introduces additional capabilities for managing data in lakehouse-style architectures.
Candidates who understand how data moves from ingestion into structured storage and processing environments will have a stronger technical foundation for this type of role.
π Engineering Practices and CI/CD
Corning expects data engineering work to follow professional software development practices. The engineer will participate in environments that use Agile development, continuous integration, continuous deployment, version control, and code reviews.
This is important because production data pipelines must be maintained over time. Changes to source systems, schemas, business requirements, or infrastructure can affect existing workflows.
Technologies and practices relevant to this environment include:
Git GitLab Jira Terraform New Relic
Candidates should understand that modern data engineering increasingly overlaps with software engineering and DevOps. Writing transformation logic is only one part of the job; maintaining reliable production workflows is equally important.
π Production Monitoring and Operations
The role also includes supporting systems in a DevOps-oriented environment. This involves monitoring jobs, tokens, and overall system performance.
Production awareness is valuable because data engineering pipelines can fail due to issues such as unavailable source systems, unexpected data formats, infrastructure problems, or processing failures.
A useful skill for candidates is learning how to investigate a failed workflow systematically:
Identify where the failure occurred.
Review logs and monitoring information.
Check the input data and source availability.
Understand the technical cause.
Apply and validate the appropriate fix.
This type of problem-solving approach is highly relevant for enterprise data engineering environments.
π€ Working Across Technical Teams
Communication is an important part of this position. The Data Engineer will interact with data source teams, domain experts, application developers, engineers, and data scientists.
Technical knowledge alone is not enough when requirements come from multiple teams. Engineers need to understand what data is required, how it should be represented, and whether the final dataset satisfies downstream requirements.
The job also mentions communicating data modeling decisions clearly to users and other technical teams.
Technical Communication Matters
Candidates should be prepared to explain technical decisions in practical language, especially when working with people who understand the business or domain but may not work directly with data engineering tools.
π What Candidates Should Learn
For candidates preparing for a role like this, the strongest learning areas come directly from the technologies and responsibilities mentioned in the position.
Focus on understanding ETL and ELT concepts, especially how data travels from source systems into data lakes and analytical environments.
Build practical familiarity with Python and SQL, then learn how distributed processing works with Apache Spark. Understanding file formats such as Parquet and modern architectures such as Delta Lake can also provide a strong foundation.
Cloud fundamentals are another useful area. Learn the purpose of services such as S3, EC2, RDS, EMR, Lambda, and CloudWatch rather than memorizing service names.
A small project that demonstrates ingestion, transformation, validation, and storage can be more useful on a resume than simply listing technologies.
π― What Recruiters May Evaluate
Candidates may benefit from being prepared to discuss the following areas:
Data pipeline fundamentals: Explain how data can move from a source system into a data lake.
Apache Spark concepts: Understand distributed processing and the basic architecture behind Spark workloads.
ETL and ELT: Know the difference between transforming data before storage and transforming it after loading into a target platform.
Cloud storage: Understand why services such as Amazon S3 are commonly used in data lake architectures.
Data quality: Be able to explain why validation and profiling are important.
Version control: Know how Git supports collaborative software development.
Problem-solving: Be prepared to explain how you would investigate a failed data pipeline.
π Resume Strengthening Areas
Candidates applying for data engineering opportunities should make their technical work easy for recruiters to understand.
Instead of listing only technology names, describe practical usage.
For example:
Built a Python and Apache Spark pipeline to ingest raw datasets, transform records, validate data quality, and store processed data in Parquet format.
This type of description communicates both the technologies used and the engineering work completed.
Projects involving AWS, Spark, Delta Lake, SQL, Python, data validation, CI/CD, or data lake architectures can be particularly relevant to this opportunity.
π Keywords for Resume
Python β’ Java β’ Scala β’ Kotlin β’ Apache Spark β’ ETL β’ ELT β’ Data Engineering β’ Data Lakes β’ AWS β’ Amazon S3 β’ EC2 β’ AWS DMS β’ RDS β’ EMR β’ Redshift β’ Lambda β’ DynamoDB β’ CloudWatch β’ CloudTrail β’ Parquet β’ Delta Lake β’ SQL β’ Data Validation β’ Data Profiling β’ Data Cleansing β’ Data Enrichment β’ Git β’ GitLab β’ CI/CD β’ Jira β’ Terraform β’ New Relic β’ JupyterHub β’ Agile Development β’ Data Modeling β’ DevOps
π‘ Final Perspective
This Corning opportunity is particularly relevant for candidates interested in enterprise data engineering and advanced analytics platforms. The role combines data ingestion, quality validation, cloud technologies, Apache Spark, Delta Lake, software engineering practices, and collaboration with technical teams. It offers a practical environment for building skills that are increasingly important in modern large-scale data platforms.
The above article is written by me, a person interested in technology, automobiles, modern gadgets, movies, music, and clean aesthetics.



