Skip to content
Technologies / Data & Analytics

Scaling Data Analysis with Python and Dask

This training is dedicated to practical use of the Dask library for scaling data analyses in Python. The workshop program is designed to allow participants to transition from standard analyses to distributed processing on large datasets. During sessions, participants will learn not only the theoretical foundations of parallel processing but primarily how to transform existing Pandas and NumPy analyses into efficient solutions utilizing Dask capabilities. Practical workshops, comprising 70% of training time, are based on real scenarios and problems encountered in daily data analyst work.

Required Participant Preparation

  • Advanced Python knowledge
  • Pandas and NumPy experience
  • Data processing fundamentals understanding
  • Parallel programming concepts knowledge

Benefits

  • Ability to transform analyses to distributed environments
  • Practical Dask library knowledge
  • Analysis performance optimization skills
  • Memory management in large computations ability
  • Distributed code debugging techniques knowledge
  • Production environment configuration experience
  • Parallel processing principles understanding

Who is this training for?

Data analysts working with large datasets
Python programmers specializing in data analysis
Data Scientists seeking performance solutions
Data engineers responsible for process optimization
Analytical solution architects
Machine Learning specialists working with large datasets
Analytical application developers

Training program

01

Dask architecture and operating principles

  • Comparison with traditional Python tools
  • Distributed environment configuration
02

Basic Dask data structures

  • Transforming Analyses to Distributed Environment
  • Migrating Pandas code to Dask DataFrame
  • Grouping and aggregation operation optimization
03

Data stream processing

  • Memory management
  • Advanced Processing Techniques
  • Matrix computations with Dask Array
  • Parallel task processing
  • Computation graph optimization
  • Debugging and profiling
  • Deploying Production Solutions
04

Dask cluster configuration

  • Monitoring and diagnostics
  • Big Data ecosystem integration
  • Scaling strategies

Delivery Methods

Online

  • Convenience of participating from anywhere
  • Interactive live sessions with trainer
  • Materials available for 30 days
  • No travel costs

On-site

  • Direct contact with trainer and group
  • Intensive hands-on workshops
  • Networking with other participants
  • Full focus on learning

Frequently asked questions

What are the prerequisites for this training?

The Scaling Data Analysis with Python and Dask training does not require specialized prior knowledge. Basic IT knowledge is sufficient.

What is the format and duration of this training?

The training lasts 2 days and is available in online and on-site format. Sessions run from 9:00 AM to 4:00 PM. We can also customize the schedule to fit your team's needs.

Who is this training designed for?

This training is designed for: Data analysts working with large datasets; Python programmers specializing in data analysis; Data Scientists seeking performance solutions.

Klaudia Janecka
Klaudia Janecka Opiekun szkolenia

Request a quote

Funding Options

Check funding options for your company

Up to 80%

Development Services Database

Up to 80% funding for SMEs from EU funds

Check availability
Up to 100%

National Training Fund

Up to 100% funding for employers

Learn more

Trusted by

We train teams at Poland's largest companies

ING Bank - EITT client
mBank - EITT client
PKO Bank Polski - EITT client
PZU - EITT client
Allianz - EITT client
T-Mobile - EITT client
KGHM - EITT client
PGE - EITT client
IKEA - EITT client
InPost - EITT client
Leroy Merlin - EITT client
ZUS - EITT client

Interested in this training?

Contact us - we'll prepare an offer tailored to your organization's needs.

500+ experts
2500+ trainings available
ISO 9001 quality certified
Request Training
Call us +48 22 487 84 90