Data preparation and quality control in Python
Practical training on data preparation and quality control in Python using pandas, Great Expectations, scikit-learn imputers, ydata-profiling and Pydantic. Participants will learn data wrangling techniques, data schema validation, missing data handling, dataset profiling, and building ETL pipelines with data versioning (DVC).
Why choose this training?
Input data quality determines the quality of every analysis and every ML model — and in practice, analysts and data scientists spend most of their time cleaning and validating data. Python with its ecosystem of pandas, Great Expectations, and scikit-learn offers powerful tools for systematic data quality management at every stage of the pipeline.
During the training, participants will learn the complete workflow — from data wrangling with pandas, through validation with Great Expectations and Pydantic, missing data handling with scikit-learn imputers, profiling with ydata-profiling, to building reproducible ETL pipelines with data versioning (DVC).
After completing the training, participants will be able to: efficiently process and transform data using pandas, define and automate validation rules with Great Expectations, apply Pydantic for data schema validation in pipelines, analyze missing data patterns and apply scikit-learn imputers. These competencies directly translate into higher efficiency in IT project execution.
This training is particularly valuable for: data analysts responsible for data quality in their organization, data scientists working with unclean datasets, data engineers building ETL pipelines in Python.
What sets our approach apart?
At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.
With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.
Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.
Benefits
- Efficiently process and transform data using pandas
- Define and automate validation rules with Great Expectations
- Apply Pydantic for data schema validation in pipelines
- Analyze missing data patterns and apply scikit-learn imputers
- Create automated profiling reports with ydata-profiling
- Detect outliers and anomalies using statistical and ML methods
- Design reproducible ETL pipelines with data versioning (DVC)
- Integrate data quality control into existing analytical workflows
Who is this training for?
Prerequisites
- Basic knowledge of Python (variables, functions, classes, modules)
- Experience working with pandas (DataFrame, Series)
- Basic familiarity with numpy and matplotlib (recommended)
- Own laptop with Python 3.10+ and a virtual environment installed
Training program
Data wrangling with pandas
- Loading data from various sources (CSV, Excel, SQL, API)
- Filtering, grouping and aggregation with pandas
- Working with text data (str accessor, regex)
- Date and time zone handling
- Combining and merging datasets (merge, join, concat)
Data validation with Great Expectations and Pydantic
- Great Expectations — defining expectations and suites
- Automatic generation of validation profiles
- Data Docs — interactive validation reports
- Pydantic — data schema validation in pipelines
- Integrating validation into processing workflows
Missing data handling and imputation
- Analyzing missing data patterns (missingno, seaborn)
- Missing data mechanisms (MCAR, MAR, MNAR)
- scikit-learn imputers (SimpleImputer, KNNImputer, IterativeImputer)
- Imputation strategies for different data types
- Imputation quality assessment and result validation
Data profiling and anomaly detection
- ydata-profiling — automated profiling reports
- Outlier detection (IQR, Z-score, Isolation Forest)
- Distribution and correlation analysis
- Duplicate and inconsistency detection
- Data quality visualization with matplotlib and seaborn
ETL pipelines and data versioning
- Designing ETL pipelines in Python (Prefect, Luigi)
- DVC — data versioning and reproducibility
- Transformation logging and change auditing
- Pipeline testing with pytest
- Best practices for organizing data quality projects
Delivery Methods
Online
- Convenience of participating from anywhere
- Interactive live sessions with trainer
- Materials available for 30 days
- No travel costs
On-site
- Direct contact with trainer and group
- Intensive hands-on workshops
- Networking with other participants
- Full focus on learning
Frequently asked questions
What prerequisites do I need to meet before the training?
Basic knowledge of Python (variables, functions, classes) and experience working with pandas (DataFrame, Series) are required. Familiarity with numpy and matplotlib is recommended but not mandatory — key elements will be reviewed during the training.
What format is the training delivered in and how long does it last?
The training lasts 2 days and is available in online (live) and on-site formats. The program is based on hands-on exercises — participants work with real, messy datasets. Each participant receives training materials and a completion certificate.
Who is this training designed for?
The training is designed for data analysts responsible for data quality, data scientists working with unclean datasets, data engineers building ETL pipelines in Python, and ML engineers ensuring input data quality for models.
Request a quote
Funding Options
Check funding options for your company
Development Services Database
Up to 80% funding for SMEs from EU funds
Check availabilityNational Training Fund
Up to 100% funding for employers
Learn moreTrusted by
We train teams at Poland's largest companies
Interested in this training?
Contact us - we'll prepare an offer tailored to your organization's needs.