Large-scale data analysis in R
Training on large-scale data analysis in R, covering memory optimization, parallel processing and integration with databases and Apache Spark. The program focuses on practical use of data.table, dplyr, future, foreach and sparklyr packages for efficient work with datasets that exceed the capabilities of standard tools. Participants will learn chunked processing techniques, memory profiling and query optimization.
Why choose this training?
Modern data analysis increasingly requires working with datasets that exceed the capabilities of standard R tools. This two-day training prepares analysts and Data Scientists for efficient processing of large data volumes — from optimizing data structures through data.table and dplyr, through parallel processing with future and foreach packages, to database integration (DBI, dbplyr) and distributed processing in Apache Spark via sparklyr. The program emphasizes practical techniques for memory management, chunked processing and choosing the right data formats.
After completing the training, participants will be able to: efficiently process large datasets using data.table and dplyr, implement parallel processing with future and foreach packages, optimize RAM usage and apply chunked processing techniques, integrate R with databases through DBI and dbplyr. These competencies directly translate into higher efficiency in IT project execution.
This training is particularly valuable for: data analysts working with large datasets, Data Scientists using R in their daily work, statisticians and researchers processing large data volumes.
What sets our approach apart?
At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.
With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.
Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.
Benefits
- Participants will be able to efficiently process large datasets in R using data.table and dplyr
- They will gain the ability to implement parallel processing with future and foreach packages
- They will learn to optimize RAM usage and apply chunked processing techniques
- They will learn methods of integrating R with databases through DBI and dbplyr
- They will be able to configure and use Apache Spark through sparklyr for distributed processing
- They will develop skills in profiling R code performance and identifying bottlenecks
- They will master techniques for working with data formats optimized for large datasets (Arrow, fst, Parquet)
Who is this training for?
Prerequisites
- Intermediate knowledge of R programming language
- Basic familiarity with tidyverse (dplyr, tidyr)
- Experience in tabular data analysis
- Basic knowledge of SQL
Training program
Efficient data structures in R
- data.table package — syntax, indexing and group operations
- Comparing data.table vs dplyr for performance
- Advanced dplyr operations on large datasets
- Memory profiling and object management in R
Parallel processing
- future framework — planning and executing parallel tasks
- foreach package with doParallel and doFuture backends
- Data chunking strategies
- Progress monitoring and error handling in parallel processing
Memory optimization and chunked processing
- Techniques for reducing RAM usage
- Processing data in chunks (chunked processing)
- Lazy evaluation and deferred computations
- Memory-efficient data formats (fst, qs, Arrow)
Database integration
- DBI package — connections to PostgreSQL, MySQL, SQLite
- dbplyr — translating dplyr code to SQL
- Query optimization and data retrieval
- Working with data directly in the database without loading into memory
Apache Spark via sparklyr
- Configuration and connecting to a Spark cluster
- Distributed processing with dplyr interface
- Spark SQL and operations on large tables
- Integrating sparklyr with analytical pipelines
Delivery Methods
Online
- Convenience of participating from anywhere
- Interactive live sessions with trainer
- Materials available for 30 days
- No travel costs
On-site
- Direct contact with trainer and group
- Intensive hands-on workshops
- Networking with other participants
- Full focus on learning
Frequently asked questions
What are the prerequisites for this training?
The Large-scale data analysis in R training requires intermediate knowledge of R programming language, basic familiarity with tidyverse (dplyr, tidyr), experience in tabular data analysis and basic knowledge of SQL.
What is the format and duration of this training?
The training lasts 2 days and is available in online and on-site format.
Who is this training designed for?
This training is designed for data analysts working with large datasets, Data Scientists using R in their daily work, and statisticians and researchers processing large data volumes.
Request a quote
Funding Options
Check funding options for your company
Development Services Database
Up to 80% funding for SMEs from EU funds
Check availabilityNational Training Fund
Up to 100% funding for employers
Learn moreTrusted by
We train teams at Poland's largest companies
Interested in this training?
Contact us - we'll prepare an offer tailored to your organization's needs.