Skip to content
Technologies

Data preparation and quality control in Python

Practical training on data preparation and quality control in Python using pandas, Great Expectations, scikit-learn imputers, ydata-profiling and Pydantic. Participants will learn data wrangling techniques, data schema validation, missing data handling, dataset profiling, and building ETL pipelines with data versioning (DVC).

Why choose this training?

Input data quality determines the quality of every analysis and every ML model — and in practice, analysts and data scientists spend most of their time cleaning and validating data. Python with its ecosystem of pandas, Great Expectations, and scikit-learn offers powerful tools for systematic data quality management at every stage of the pipeline.

During the training, participants will learn the complete workflow — from data wrangling with pandas, through validation with Great Expectations and Pydantic, missing data handling with scikit-learn imputers, profiling with ydata-profiling, to building reproducible ETL pipelines with data versioning (DVC).

After completing the training, participants will be able to: efficiently process and transform data using pandas, define and automate validation rules with Great Expectations, apply Pydantic for data schema validation in pipelines, analyze missing data patterns and apply scikit-learn imputers. These competencies directly translate into higher efficiency in IT project execution.

This training is particularly valuable for: data analysts responsible for data quality in their organization, data scientists working with unclean datasets, data engineers building ETL pipelines in Python.

What sets our approach apart?

At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.

With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.

Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.

Benefits

  • Efficiently process and transform data using pandas
  • Define and automate validation rules with Great Expectations
  • Apply Pydantic for data schema validation in pipelines
  • Analyze missing data patterns and apply scikit-learn imputers
  • Create automated profiling reports with ydata-profiling
  • Detect outliers and anomalies using statistical and ML methods
  • Design reproducible ETL pipelines with data versioning (DVC)
  • Integrate data quality control into existing analytical workflows

Who is this training for?

Data analysts responsible for data quality in their organization
Data Scientists working with unclean datasets
Data Engineers building ETL pipelines in Python
Python developers expanding competencies into data engineering
Business Intelligence specialists
ML Engineers ensuring input data quality for models
Data managers and data governance professionals

Prerequisites

  • Basic knowledge of Python (variables, functions, classes, modules)
  • Experience working with pandas (DataFrame, Series)
  • Basic familiarity with numpy and matplotlib (recommended)
  • Own laptop with Python 3.10+ and a virtual environment installed

Training program

01

Data wrangling with pandas

  • Loading data from various sources (CSV, Excel, SQL, API)
  • Filtering, grouping and aggregation with pandas
  • Working with text data (str accessor, regex)
  • Date and time zone handling
  • Combining and merging datasets (merge, join, concat)
02

Data validation with Great Expectations and Pydantic

  • Great Expectations — defining expectations and suites
  • Automatic generation of validation profiles
  • Data Docs — interactive validation reports
  • Pydantic — data schema validation in pipelines
  • Integrating validation into processing workflows
03

Missing data handling and imputation

  • Analyzing missing data patterns (missingno, seaborn)
  • Missing data mechanisms (MCAR, MAR, MNAR)
  • scikit-learn imputers (SimpleImputer, KNNImputer, IterativeImputer)
  • Imputation strategies for different data types
  • Imputation quality assessment and result validation
04

Data profiling and anomaly detection

  • ydata-profiling — automated profiling reports
  • Outlier detection (IQR, Z-score, Isolation Forest)
  • Distribution and correlation analysis
  • Duplicate and inconsistency detection
  • Data quality visualization with matplotlib and seaborn
05

ETL pipelines and data versioning

  • Designing ETL pipelines in Python (Prefect, Luigi)
  • DVC — data versioning and reproducibility
  • Transformation logging and change auditing
  • Pipeline testing with pytest
  • Best practices for organizing data quality projects

Delivery Methods

Online

  • Convenience of participating from anywhere
  • Interactive live sessions with trainer
  • Materials available for 30 days
  • No travel costs

On-site

  • Direct contact with trainer and group
  • Intensive hands-on workshops
  • Networking with other participants
  • Full focus on learning

Frequently asked questions

What prerequisites do I need to meet before the training?

Basic knowledge of Python (variables, functions, classes) and experience working with pandas (DataFrame, Series) are required. Familiarity with numpy and matplotlib is recommended but not mandatory — key elements will be reviewed during the training.

What format is the training delivered in and how long does it last?

The training lasts 2 days and is available in online (live) and on-site formats. The program is based on hands-on exercises — participants work with real, messy datasets. Each participant receives training materials and a completion certificate.

Who is this training designed for?

The training is designed for data analysts responsible for data quality, data scientists working with unclean datasets, data engineers building ETL pipelines in Python, and ML engineers ensuring input data quality for models.

Anna Polak
Anna Polak Opiekun szkolenia

Request a quote

Funding Options

Check funding options for your company

Up to 80%

Development Services Database

Up to 80% funding for SMEs from EU funds

Check availability
Up to 100%

National Training Fund

Up to 100% funding for employers

Learn more

Trusted by

We train teams at Poland's largest companies

ING Bank - EITT client
mBank - EITT client
PKO Bank Polski - EITT client
PZU - EITT client
Allianz - EITT client
T-Mobile - EITT client
KGHM - EITT client
PGE - EITT client
IKEA - EITT client
InPost - EITT client
Leroy Merlin - EITT client
ZUS - EITT client

Interested in this training?

Contact us - we'll prepare an offer tailored to your organization's needs.

500+ experts
2500+ trainings available
ISO 9001 quality certified
Request Training
Call us +48 22 487 84 90