Skip to content
Technologies

Data preparation and quality control in R

Practical training on data preparation and quality control in R using the tidyverse ecosystem. Participants will learn the data processing pipeline (dplyr, tidyr, stringr, lubridate), validation tools (pointblank, validate), missing data handling methods (mice, naniar), outlier detection, data profiling, and building reproducible ETL workflows.

Why choose this training?

Data quality is the foundation of every reliable analysis — and in practice, analysts spend up to 80% of their time cleaning and preparing data. Without a systematic approach to validation, hidden errors propagate through the entire analytical pipeline. R with the tidyverse ecosystem offers some of the best tools for structural data processing and quality control.

During the training, participants will learn the complete pipeline — from loading raw data, through transformation with tidyverse, validation with pointblank and validate, missing data handling with mice and naniar, to building reproducible ETL workflows with the targets package.

After completing the training, participants will be able to: build data processing pipelines using tidyverse, define and automate validation rules with pointblank and validate, analyze missing data patterns and apply appropriate imputation methods, detect outliers using statistical methods and visualize anomalies. These competencies directly translate into higher efficiency in IT project execution.

This training is particularly valuable for: data analysts responsible for data quality in their organization, data scientists working with unclean datasets, data engineers building ETL pipelines in R.

What sets our approach apart?

At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.

With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.

Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.

Benefits

  • Build data processing pipelines using tidyverse (dplyr, tidyr, stringr, lubridate)
  • Define and automate data validation rules with pointblank and validate
  • Analyze missing data patterns and apply appropriate imputation methods (mice)
  • Detect outliers using statistical methods and visualize anomalies
  • Create automated data profiling and quality reports
  • Design reproducible ETL workflows with the targets package
  • Integrate data quality control into existing analytical pipelines

Who is this training for?

Data analysts responsible for data quality in their organization
Data Scientists working with unclean datasets
Data Engineers building ETL pipelines in R
Statisticians and researchers ensuring analysis reproducibility
Business Intelligence specialists
R programmers looking to standardize data cleaning processes
Data managers and data governance professionals

Prerequisites

  • Basic knowledge of R (variables, functions, data structures)
  • Experience working with data frames (data.frame / tibble)
  • Basic familiarity with tidyverse (dplyr, ggplot2) is recommended
  • Own laptop with R and RStudio installed

Training program

01

Data processing pipeline with tidyverse

  • Tidyverse architecture and tidy data philosophy
  • dplyr — filtering, transformation and data aggregation
  • tidyr — pivoting, separating and combining columns
  • stringr — text data processing and regex
  • lubridate — date and time zone handling
02

Data validation and quality control

  • pointblank — declarative validation with HTML reports
  • validate — validation rules and quality monitoring
  • Defining business rules and acceptance thresholds
  • Automated data quality reports
  • Integrating validation into processing pipelines
03

Missing data handling and imputation

  • naniar — visualization and analysis of missing data patterns
  • Missing data mechanisms (MCAR, MAR, MNAR)
  • mice — multiple imputation by chained equations
  • Comparing imputation methods and evaluating their quality
  • Strategies for handling missing values in different data types
04

Outlier detection and data profiling

  • Statistical methods for outlier detection (IQR, Z-score, Mahalanobis)
  • Data profiling packages (skimr, DataExplorer)
  • Automated dataset profiling reports
  • Distribution and anomaly visualization
  • Strategies for handling outliers
05

Reproducible ETL workflows in R

  • Designing ETL pipelines with the targets package
  • Data cleaning as a reproducible workflow
  • Logging and auditing data transformations
  • Pipeline testing with testthat
  • Best practices for organizing data quality projects

Delivery Methods

Online

  • Convenience of participating from anywhere
  • Interactive live sessions with trainer
  • Materials available for 30 days
  • No travel costs

On-site

  • Direct contact with trainer and group
  • Intensive hands-on workshops
  • Networking with other participants
  • Full focus on learning

Frequently asked questions

What prerequisites do I need to meet before the training?

Basic knowledge of R (variables, functions, data structures) and experience working with data frames are required. Familiarity with tidyverse is recommended but not mandatory — key functions will be reviewed at the beginning of the training.

What format is the training delivered in and how long does it last?

The training lasts 2 days and is available in online (live) and on-site formats. The program is based on hands-on exercises — participants work with real, messy datasets. Each participant receives training materials and a completion certificate.

Who is this training designed for?

The training is designed for data analysts responsible for data quality, data scientists working with unclean datasets, data engineers building ETL pipelines in R, and statisticians ensuring analysis reproducibility.

Anna Polak
Anna Polak Opiekun szkolenia

Request a quote

Funding Options

Check funding options for your company

Up to 80%

Development Services Database

Up to 80% funding for SMEs from EU funds

Check availability
Up to 100%

National Training Fund

Up to 100% funding for employers

Learn more

Trusted by

We train teams at Poland's largest companies

ING Bank - EITT client
mBank - EITT client
PKO Bank Polski - EITT client
PZU - EITT client
Allianz - EITT client
T-Mobile - EITT client
KGHM - EITT client
PGE - EITT client
IKEA - EITT client
InPost - EITT client
Leroy Merlin - EITT client
ZUS - EITT client

Interested in this training?

Contact us - we'll prepare an offer tailored to your organization's needs.

500+ experts
2500+ trainings available
ISO 9001 quality certified
Request Training
Call us +48 22 487 84 90