Data preparation and quality control in R
Practical training on data preparation and quality control in R using the tidyverse ecosystem. Participants will learn the data processing pipeline (dplyr, tidyr, stringr, lubridate), validation tools (pointblank, validate), missing data handling methods (mice, naniar), outlier detection, data profiling, and building reproducible ETL workflows.
Why choose this training?
Data quality is the foundation of every reliable analysis — and in practice, analysts spend up to 80% of their time cleaning and preparing data. Without a systematic approach to validation, hidden errors propagate through the entire analytical pipeline. R with the tidyverse ecosystem offers some of the best tools for structural data processing and quality control.
During the training, participants will learn the complete pipeline — from loading raw data, through transformation with tidyverse, validation with pointblank and validate, missing data handling with mice and naniar, to building reproducible ETL workflows with the targets package.
After completing the training, participants will be able to: build data processing pipelines using tidyverse, define and automate validation rules with pointblank and validate, analyze missing data patterns and apply appropriate imputation methods, detect outliers using statistical methods and visualize anomalies. These competencies directly translate into higher efficiency in IT project execution.
This training is particularly valuable for: data analysts responsible for data quality in their organization, data scientists working with unclean datasets, data engineers building ETL pipelines in R.
What sets our approach apart?
At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.
With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.
Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.
Benefits
- Build data processing pipelines using tidyverse (dplyr, tidyr, stringr, lubridate)
- Define and automate data validation rules with pointblank and validate
- Analyze missing data patterns and apply appropriate imputation methods (mice)
- Detect outliers using statistical methods and visualize anomalies
- Create automated data profiling and quality reports
- Design reproducible ETL workflows with the targets package
- Integrate data quality control into existing analytical pipelines
Who is this training for?
Prerequisites
- Basic knowledge of R (variables, functions, data structures)
- Experience working with data frames (data.frame / tibble)
- Basic familiarity with tidyverse (dplyr, ggplot2) is recommended
- Own laptop with R and RStudio installed
Training program
Data processing pipeline with tidyverse
- Tidyverse architecture and tidy data philosophy
- dplyr — filtering, transformation and data aggregation
- tidyr — pivoting, separating and combining columns
- stringr — text data processing and regex
- lubridate — date and time zone handling
Data validation and quality control
- pointblank — declarative validation with HTML reports
- validate — validation rules and quality monitoring
- Defining business rules and acceptance thresholds
- Automated data quality reports
- Integrating validation into processing pipelines
Missing data handling and imputation
- naniar — visualization and analysis of missing data patterns
- Missing data mechanisms (MCAR, MAR, MNAR)
- mice — multiple imputation by chained equations
- Comparing imputation methods and evaluating their quality
- Strategies for handling missing values in different data types
Outlier detection and data profiling
- Statistical methods for outlier detection (IQR, Z-score, Mahalanobis)
- Data profiling packages (skimr, DataExplorer)
- Automated dataset profiling reports
- Distribution and anomaly visualization
- Strategies for handling outliers
Reproducible ETL workflows in R
- Designing ETL pipelines with the targets package
- Data cleaning as a reproducible workflow
- Logging and auditing data transformations
- Pipeline testing with testthat
- Best practices for organizing data quality projects
Delivery Methods
Online
- Convenience of participating from anywhere
- Interactive live sessions with trainer
- Materials available for 30 days
- No travel costs
On-site
- Direct contact with trainer and group
- Intensive hands-on workshops
- Networking with other participants
- Full focus on learning
Frequently asked questions
What prerequisites do I need to meet before the training?
Basic knowledge of R (variables, functions, data structures) and experience working with data frames are required. Familiarity with tidyverse is recommended but not mandatory — key functions will be reviewed at the beginning of the training.
What format is the training delivered in and how long does it last?
The training lasts 2 days and is available in online (live) and on-site formats. The program is based on hands-on exercises — participants work with real, messy datasets. Each participant receives training materials and a completion certificate.
Who is this training designed for?
The training is designed for data analysts responsible for data quality, data scientists working with unclean datasets, data engineers building ETL pipelines in R, and statisticians ensuring analysis reproducibility.
Request a quote
Funding Options
Check funding options for your company
Development Services Database
Up to 80% funding for SMEs from EU funds
Check availabilityNational Training Fund
Up to 100% funding for employers
Learn moreTrusted by
We train teams at Poland's largest companies
Interested in this training?
Contact us - we'll prepare an offer tailored to your organization's needs.