Large-scale data analysis in Python
Training on large-scale data analysis in Python, covering pandas optimization, distributed processing with Dask and PySpark, and modern libraries Polars and Vaex. The program focuses on practical techniques for memory profiling, chunked processing, working with Parquet and Arrow formats, and database integration. Participants will learn strategies for choosing the right tools depending on data scale and characteristics.
Why choose this training?
Standard Python data analysis tools quickly hit limitations when working with large datasets — pandas slows down, RAM runs out, and processing takes hours instead of minutes. This two-day training prepares analysts and data engineers for efficient work with large data volumes, from pandas optimization through modern libraries Polars and Vaex, to distributed processing with Dask and PySpark. The program also covers columnar formats Parquet and Arrow, memory profiling and strategies for choosing the right tools depending on data scale.
After completing the training, participants will be able to: optimize pandas for performance and memory usage, process distributed datasets using Dask, use the Polars library as an efficient alternative to pandas, work with columnar formats Parquet and Arrow. These competencies directly translate into higher efficiency in IT project execution.
This training is particularly valuable for: data analysts working with large datasets in Python, Data Scientists looking for more efficient data processing tools, data engineers building analytical pipelines.
What sets our approach apart?
At EITT, we believe the best learning happens through practice. During 2 days of intensive training, participants work on real-world examples and scenarios, ensuring not only theoretical understanding but above all the ability to apply it in practice.
With over 2,500 trainings in our portfolio and a 4.8/5 rating from participants, EITT is a trusted partner in competency development for organizations of all sizes. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.
Looking for training tailored to your team’s needs? Contact us — we’ll prepare a program customized to your requirements.
Benefits
- Participants will be able to optimize pandas for performance and memory usage
- They will gain the ability to process distributed datasets using Dask
- They will learn to use the Polars library as an efficient alternative to pandas
- They will learn large-scale data processing techniques with PySpark
- They will be able to work with columnar formats Parquet and Arrow
- They will develop skills in memory profiling and identifying performance bottlenecks
- They will master chunked processing and out-of-core processing techniques
- They will learn to choose the right tool depending on data scale and characteristics
Who is this training for?
Prerequisites
- Intermediate knowledge of Python
- Basic familiarity with the pandas library
- Experience in tabular data analysis
- Basic knowledge of SQL
Training program
Pandas optimization
- Memory profiling and bottleneck identification
- Data type optimization (category, int8/16/32, nullable dtypes)
- Operation vectorization and avoiding loops
- Chunked reading with read_csv and read_parquet
Dask and distributed processing
- Dask architecture — DataFrame, Array, Bag
- Lazy evaluation and computation graph
- Local and distributed scheduler configuration
- Dask diagnostics and dashboard
Polars and Vaex
- Polars — Rust architecture, lazy and eager mode
- Polars expressions vs pandas syntax
- Vaex — out-of-core processing on large datasets
- Performance comparison of Polars vs pandas vs Dask
PySpark and data formats
- PySpark DataFrame API — operations on large tables
- Columnar formats Parquet and Arrow (PyArrow)
- Conversion between pandas, Polars and Spark DataFrame
- Partitioning and data compression optimization
Database integration and pipelines
- SQLAlchemy and connectors for relational databases
- Chunked processing with databases
- Building ETL pipelines using the tools covered
- Strategies for tool selection depending on data scale
Delivery Methods
Online
- Convenience of participating from anywhere
- Interactive live sessions with trainer
- Materials available for 30 days
- No travel costs
On-site
- Direct contact with trainer and group
- Intensive hands-on workshops
- Networking with other participants
- Full focus on learning
Frequently asked questions
What are the prerequisites for this training?
The Large-scale data analysis in Python training requires intermediate knowledge of Python, basic familiarity with the pandas library, experience in tabular data analysis and basic knowledge of SQL.
What is the format and duration of this training?
The training lasts 2 days and is available in online and on-site format.
Who is this training designed for?
This training is designed for data analysts working with large datasets in Python, Data Scientists looking for more efficient data processing tools, and data engineers building analytical pipelines.
Request a quote
Funding Options
Check funding options for your company
Development Services Database
Up to 80% funding for SMEs from EU funds
Check availabilityNational Training Fund
Up to 100% funding for employers
Learn moreTrusted by
We train teams at Poland's largest companies
Interested in this training?
Contact us - we'll prepare an offer tailored to your organization's needs.