Big Data Training
Big Data training: Hadoop, Spark, Kafka, Hive, Flink. Practical courses in large-scale data processing and analytics.
84 courses available
Beginner
Fundamentals of machine learning with Scala and Apache Spark
The training introduces participants to the world of big data processing and machine learning using Scala and Apache Spark. The program focuses on the practical use of MLlib, Spark's machine learning library, to build scalable analytics solutions. The classes are conducted in the form of workshops, where participants work on real data sets, implementing a variety of ML algorithms.
from 2,450 PLNHadoop for Developers - Four-Day Course
This training provides developers with a thorough introduction to building applications in the Apache Hadoop ecosystem. The program covers all key aspects of Big Data programming, from MapReduce basics to advanced data processing techniques. Practical workshops constitute the main element of the training, allowing participants to independently implement solutions in various Hadoop ecosystem technologies. Classes emphasize practical applications and real project scenarios.
from 4,500 PLNApache Spark basics - from theory to practice
The training provides a thorough knowledge of Apache Spark fundamentals, combining theoretical foundations with practical application. The program covers key aspects of data processing, from basic operations to advanced transformations. Hands-on workshops allow participants to gain hands-on experience in designing and implementing Spark-based solutions.
from 3,900 PLNHadoop Administration - Fundamentals and Best Practices
The training introduces participants to the world of Apache Hadoop platform administration, focusing on fundamental aspects of Big Data infrastructure management. The program covers key elements of the Hadoop ecosystem, presenting them in the context of practical administrative tasks. Classes combine theory with intensive practical workshops, during which participants gain experience in configuring, monitoring, and maintaining Hadoop clusters. The training methodology is based on gradually building competencies through practical exercises and real-life scenarios.
from 3,300 PLNApache Flink basics
Intensive introductory training on Apache Flink technology, focusing on the foundations of stream processing and practical applications of the platform. The program covers a holistic approach to building streaming applications, from basic concepts to advanced implementation patterns. The workshop combines theory with intensive hands-on experience, enabling participants to gain practical skills in designing, implementing and deploying Apache Flink-based solutions. The training uses real use cases and best practices from production projects.
from 4,500 PLNBasics of Big Data and database systems
The training introduces participants to the foundations of Big Data systems and modern databases, focusing on the practical aspects of their use in a production environment. The program combines theory with practice through workshops and exercises on real-world examples. Participants will learn key concepts, architectures and tools used in the Big Data ecosystem. Classes are conducted in an interactive format, enabling immediate application of the knowledge gained.
from 2,450 PLNBig Data and Data Science
This training provides intensive introduction to the world of Data Science in the Big Data context, combining theoretical fundamentals with practical application. The program is structured so that participants can understand key Data Science concepts and their implementation in large dataset environments. Sessions are conducted in workshop format, where each theoretical topic is immediately verified through practical exercises on real datasets. The teaching methodology focuses on building comprehensive understanding of the data analysis process, from initial exploration to advanced modeling.
from 2,450 PLNBIG DATA - data science (Basic Level)
Gain basic skills in data analysis and Data Science. Training includes an introduction to the R language, data operations, and creating visualizations and interactive reports.
from 3,100 PLNPractical Introduction to Data Analysis and Big Data
Practical introduction to data analysis and Big Data concepts. Learn fundamental tools, techniques, and methodologies for modern analytics.
from 2,450 PLNA practical introduction to stream processing
The training focuses on the practical aspects of implementing stream processing systems in a Big Data environment. The program combines theory with intensive hands-on workshops where participants learn about architecture, design patterns and best practices for stream processing. The workshops are conducted in the form of hands-on labs using real use cases. The classes are based on the implementation of specific business scenarios.
from 3,900 PLNApache Accumulo basics
The training introduces participants to the world of distributed data processing using Apache Accumulo. The workshop program guides through the fundamental concepts of the Accumulo database, its architecture and practical applications in a Big Data environment. The class combines theory with intensive hands-on workshops where participants learn how to design, implement and optimize Accumulo-based solutions. All issues are presented on real business examples.
from 3,900 PLNData Architecture Fundamentals
A fundamental training in designing and implementing modern data architectures. Participants learn design methodologies, architectural patterns, and best practices for building scalable big data solutions. The program combines theory with practical workshops, enabling understanding of both concepts and their practical application in real business scenarios.
from 3,500 PLNIBM Flash Storage Fundamentals
Enterprise data is growing exponentially. This data explosion is rendering commonly accepted data management practices inadequate. Social, mobile, cloud, big data and analytics are driving the explosion of data volumes and creating new challenges for data protection, disaster recovery, regulatory requirements and compliance standards. Data management has become one of the highest priorities for today's organizations.
from 3,500 PLNIntroduction to graph processing
The training provides a thorough introduction to the concepts and practical applications of graph processing in the context of Big Data. The program combines the theoretical foundations of graph theory with practical applications in real business scenarios. Participants, through hands-on workshops, will learn methods for designing, implementing and optimizing graph-based solutions, with a focus on system performance and scalability.
from 4,500 PLNIntroduction to Storage
Data is growing at massive scale faster than ever before as new data is created every second by big data analytics, mobile devices, and social applications. At the rate data is growing, businesses of all sizes - large and small - face challenges such as capacity growth, performance impact, and day-to-day management to maintain and preserve data accuracy in data centers or in the cloud. To meet the needs of businesses and innovative devices, IBM offers robust and cognitive storage system solutions designed to analyze massive amounts of data. IBM Storage solutions provide speed and efficiency of ready data access with the agility and cost-effectiveness of hybrid cloud and software-defined storage. More importantly, product solutions that connect data in any architecture, helping businesses meet challenges while reducing total cost of ownership. Introduction to Storage, SS01, is a foundational course for the IBM Storage product portfolio. This course will cover the fundamental elements of storage concepts and storage solutions, RAID technology, and connectivity. This course also covers the concept of high availability and business continuity planning for disaster recovery, as well as discusses the concept of storage virtualization and its many benefits. Additionally, this course will highlight IBM Storage system solutions including disk, flash, and tape systems designed for all types of business needs. We also cover the SAN family to help maintain continuity, and the IBM Spectrum Suite storage management software package covers storage solutions for the cloud.
from 10,300 PLNIntermediate
Spark Streaming with Python and Kafka
The training is devoted to the implementation of real-time data stream processing systems using Apache Spark Streaming, Python and Apache Kafka. The program guides participants through the process of building scalable solutions for analyzing streaming data, combining theory with practical implementation workshops. The class focuses on real-world scenarios of using streaming technologies in Big Data projects.
from 1,650 PLNApache Spark for .NET developers
The training focuses on the use of Apache Spark in .NET for Big Data processing. The program is delivered through hands-on workshops where participants implement Big Data solutions using .NET for Apache Spark. The class combines theory and practice, allowing participants to understand and apply distributed data processing techniques.
from 2,150 PLNBig Data BI for Telecommunications Service Providers
Advanced training dedicated to using Big Data technologies in Business Intelligence context for the telecommunications sector. The program was created with specific challenges of the telecommunications industry in mind, where real-time analysis of huge data volumes plays a key role. Participants will learn advanced data processing and analysis techniques, learn to design scalable BI architecture, and implement solutions supporting business decision-making.
from 5,050 PLNHadoop and Spark for Administrators
Advanced training for Big Data system administrators, focusing on practical aspects of Hadoop and Spark cluster management. The program combines theoretical knowledge with intensive practical workshops where participants learn configuration, optimization, and maintenance of production environments. The training provides deep understanding of distributed systems architecture and develops skills necessary for effective Big Data infrastructure management.
from 6,350 PLNApache Spark SQL - structured data processing
An intensive training course dedicated to the effective use of Apache Spark SQL in processing structured data in a distributed environment. The program covers both the theoretical basics of DataFrame processing and the practical aspects of implementing efficient transformations and analysis. Participants work with real data sets, learning how to optimize queries and effectively use the Catalyst engine. The workshop is conducted in the form of practical classes, where each concept is immediately verified through the implementation of specific use cases.
from 1,850 PLNMagellan: Geospatial Analytics in Apache Spark
Specialized training in geospatial data analysis using the Magellan library in the Apache Spark environment. The training program combines fundamental concepts of spatial analysis with practical aspects of geographic data processing in a distributed environment. Participants work on real geospatial datasets, learning to implement complex analysis and visualization. The workshop is conducted as an intensive hands-on class, where theory is immediately translated into working analytical solutions. The training methodology focuses on understanding the specifics of spatial data and effectively using the platform's capabilities to analyze it.
from 2,450 PLNApache Ambari - Effective Hadoop Cluster Management
This training focuses on practical aspects of managing Hadoop clusters using the Apache Ambari platform. Participants will learn advanced configuration, monitoring, and Big Data infrastructure maintenance techniques. The program covers both theoretical foundations of Ambari architecture and practical administrative scenarios. Workshops were designed so that participants can independently perform cluster operations, from installation, through configuration, to resolving typical operational problems.
from 3,900 PLNData Analysis with Hive/HiveQL
A one-day training focusing on practical use of Apache Hive and HiveQL language for data analysis in the Hadoop environment. The workshop program covers schema design, query optimization, and integration with other analytical tools. Participants work on real datasets, learning effective use of data warehousing in the Hadoop ecosystem. The training combines theory with intensive practical exercises.
from 1,850 PLNHadoop for Business Analysts
Practical Hadoop ecosystem training oriented towards business analyst needs. The course focuses on using Hadoop tools for business data analysis without delving into technical implementation details. Participants will learn how to effectively use the platform for processing and analyzing large datasets.
from 3,900 PLNHadoop for Project Managers
This training was designed specifically for people managing Big Data projects, focusing on business and organizational aspects of Hadoop implementations. The program covers key technical issues presented from a management perspective, enabling effective decision-making and project planning. Classes combine theory with case analysis and practical planning workshops, allowing participants to understand Big Data project specifics and develop effective implementation strategies.
from 2,450 PLNHadoop with Python - Comprehensive Course
Intensive training combining Hadoop capabilities with Python power in the Big Data context. The program covers practical MapReduce applications, using Python libraries for data analysis, and integration with the Hadoop ecosystem. Workshops are conducted in the form of practical tasks with real datasets. Participants will learn to create efficient analytical solutions using Python programming best practices.
from 4,500 PLNApache Drill - SQL on Big Data
This training provides practical introduction to Apache Drill, a SQL query system for various big data sources. Participants will learn advanced data analysis techniques using familiar SQL syntax on heterogeneous data sources. The program combines theory with intensive workshops, where participants learn to design and optimize queries for distributed systems, working on real-world use scenarios.
from 3,900 PLNApache Hadoop - Data Manipulation and Transformation
This training delves into practical aspects of data processing and transformation in the Apache Hadoop ecosystem. The program is designed so that participants understand not only technical aspects of data manipulation but also learn principles of designing effective processing workflows. Practical workshops constitute a significant part of the sessions, during which participants independently implement solutions based on real-world use cases. The teaching methodology is based on gradually introducing increasingly advanced concepts, always in the context of practical applications.
from 3,900 PLNApache Kafka Connect - integration of data sources
Specialized training on using Apache Kafka Connect to integrate diverse data sources. The program covers the design, implementation and management of connectors to external systems, with an emphasis on scalability and reliability of data flow. Hands-on workshops allow participants to gain experience in configuring and monitoring different types of connectors and solving common integration problems. The training uses real-world scenarios to demonstrate best practices in data integration.
from 1,850 PLNApache Kafka for developers - architecture and implementation
Intensive workshop training devoted to the architecture and implementation of solutions based on Apache Kafka. During the course, participants will learn both the theoretical basics of the platform and the practical aspects of its use in a production environment. The training is carried out in the form of workshops, where 70% of the time is devoted to practical exercises. The classes are based on real use cases and project scenarios.
from 3,900 PLNApache Kafka for Python Developers
The training introduces participants to the world of stream data processing using Apache Kafka and Python. The program focuses on practical aspects of implementing real-time event processing systems, leading participants from Kafka architecture basics to advanced integration patterns. Sessions are implemented in workshop format, where theory immediately translates into practical implementations of real business scenarios.
from 1,850 PLNBig Data Analytics for Telecommunications Regulators
This training presents specialized application of Big Data analytics in the context of telecommunications market regulation. The program was designed with specific requirements and challenges facing telecommunications market regulators in mind. Participants will learn advanced techniques for analyzing telecommunications data through practical workshops and case studies. Sessions focus on using analytical tools for market monitoring, detecting irregularities, and supporting decision-making processes in the regulatory area.
from 2,450 PLNBig Data Analytics in Healthcare
Advanced training combining knowledge of Big Data analytics with practical applications in the healthcare sector. The program covers methods for analyzing medical data, techniques for processing clinical information, and using machine learning in diagnostics. Workshops focus on practical aspects of implementing analytical solutions considering the specifics of medical data and security requirements.
from 3,900 PLNBig Data and the management process
The training presents a holistic approach to project and process management in the context of Big Data. The program focuses on the organizational, strategic and operational aspects of implementing Big Data solutions in an organization. Participants will learn methods for effective management of analytical projects and principles for building a data-driven strategy. The classes use case studies and practical workshops to understand the challenges of transforming an organization into a data-driven one.
from 2,450 PLNBig Data Architect
The training prepares participants for the role of Big Data solution architect, focusing on the design of scalable and efficient data processing systems. The program covers both technical and business aspects of Big Data architecture. The class is conducted in the form of workshops, where participants design and implement a variety of architectural solutions. The course emphasizes the practical application of the latest technologies and design patterns.
from 6,350 PLNBIG DATA - data science (Intermediate level)
Master Big Data and Data Science analysis techniques at an intermediate level. Training includes data analysis, visualizations, predictive models and interactive reports using the R language.
from 3,200 PLNBig Data Programming in R
Training on Big Data programming with R. Learn efficient data processing using data.table, parallel processing, and distributed computing for handling large datasets.
from 2,950 PLNBig Data Storage Solutions: NoSQL
Comprehensive training on Big Data storage solutions using NoSQL databases. Learn document stores, key-value stores, columnar databases, and graph databases for modern data architectures.
from 2,950 PLNBusiness Intelligence with Big Data for Government Institutions
This training presents advanced methods of using Business Intelligence in the Big Data context for the public sector. The program combines theory with practice, focusing on specific requirements of government institutions regarding data analysis. Practical workshops allow participants to apply learned techniques in real-world scenarios. Sessions cover security aspects, regulatory compliance, and effective use of public resources.
from 5,050 PLNEvent-Driven Architecture - Apache Kafka, CQRS, and Real-Time Business
This training focuses on designing real-time event-responsive systems that provide organizations with competitive advantage. The program covers practical use of Apache Kafka, complex event processing, and real-time business analytics production. Participants will learn practical aspects of implementing event-driven architectural patterns. The architecture-first methodology ensures solid foundations for building scalable event-based systems.
from 5,400 PLNFrom Data to Decisions with Big Data and Predictive Analytics
This training focuses on the practical use of Big Data and predictive analytics in decision-making processes. The program combines statistical theory with practical aspects of processing large datasets. Through practical workshops, participants learn to transform raw data into valuable business insights. Classes are conducted using real use cases and modern analytical tools.
from 3,900 PLNHadoop Administration on MapR Platform
The training is dedicated to advanced MapR platform administration in the context of Hadoop ecosystem management. The program has been designed to provide deep understanding of MapR architecture and practical aspects of production cluster administration. Classes combine theoretical foundations with intensive practical workshops, during which participants independently perform administrative tasks in the MapR environment. The training places particular emphasis on performance, security, and reliability aspects of production systems.
from 4,500 PLNHadoop for Administrators - Cluster Management
This training provides system administrators with advanced knowledge necessary for effective Hadoop cluster management. The program covers all key aspects of administration, from infrastructure planning to advanced optimization and troubleshooting techniques. Practical workshops constitute the main element of the training, during which participants work on real Hadoop clusters. Classes are conducted based on real-life scenarios, with emphasis on practical aspects of administration.
from 3,900 PLNMachine Learning and Big Data
The training presents the practical application of Machine Learning techniques in the context of Big Data. The program is designed to help participants understand how to effectively combine Machine Learning capabilities with Big Data challenges. During the intensive workshop, participants work on real data sets, learning to implement scalable ML solutions. The class focuses on the practical aspects of implementing machine learning models in a production environment, taking into account performance and scalability.
from 1,850 PLNMicrosoft Azure Big Data Analytics Solutions
55224A-1 is a two-day instructor-led course designed for data professionals who want to expand their knowledge of creating big data analytics solutions in Microsoft Azure. Students will learn how to design solutions for batch and real-time data processing. Various methods of using Azure will be discussed and practiced in lab exercises such as Azure CLI, Azure PowerShell and Azure Portal. The 55224A-1 labs and exercises cover the first two objectives of the 70-475 exam (Designing batch, interactive and real-time solutions for Big Data). The other two objectives (Designing machine learning and cloud analytics solutions) are covered in 55224A-2.
from 1,399 PLNMySQL to Hadoop Data Migration Using Sqoop
This training focuses on practical aspects of data migration from MySQL relational databases to Hadoop environments using Sqoop. Participants will learn advanced data import and export techniques, migration process optimization, and integration with other Hadoop ecosystem components. Workshops are conducted as practical exercises on real-world use cases. The program is designed so that participants can independently plan and conduct data migration processes.
from 2,450 PLNScaling Data Pipelines with Spark NLP
This training focuses on practical aspects of scaling natural language processing pipelines using Spark NLP. The program covers advanced optimization techniques, large-scale text processing methods, and best practices for NLP solution implementation.
from 2,450 PLNSqoop and Flume in the Big Data Ecosystem
This training focuses on Sqoop and Flume tools, key components of the Big Data ecosystem for efficient data transfer and integration. The program presents practical applications of both tools in real business scenarios. Through practical workshops, participants learn to configure and use these tools for building efficient data pipelines. Sessions cover both basic functionality and advanced data flow optimization techniques.
from 1,850 PLNBuilding Effective Collaboration Based on Lumina Spark Personality Types
This training is an intensive development workshop designed for leaders and their teams, aimed at establishing a new standard of collaborative work. Using the Lumina Spark model, participants jointly discover the dynamics taking place within their group, which enables immediate improvement of communication and mutual alignment of working styles. The program focuses on building authentic relationships and understanding the diverse traits of each team member, eliminating barriers and misunderstandings that block daily effectiveness. Through joint participation in the workshop, the leader and employees gain a coherent communication language and a practical 'user manual' for their team, which sustainably improves work comfort and allows the full potential of the whole group to be used.
from 2,200 PLNSecurity in Apache Kafka
A specialized training course dedicated to the comprehensive security of the Apache Kafka platform in a production environment. The training program guides you through all key aspects of security, from authentication and authorization, to data encryption, to security auditing and monitoring. During intensive hands-on workshops, participants implement security mechanisms in various scenarios, learning to identify and mitigate potential threats. The teaching methodology combines security theory with the practical aspects of its implementation in the context of message streaming systems.
from 1,850 PLNApache Kylin - From Classic OLAP to Real-Time Data Warehouse
This training introduces participants to the world of Apache Kylin, an advanced OLAP engine for big data. The program focuses on practical use of the platform for building efficient real-time analytical solutions. Participants will learn system architecture and how to design and implement analytical solutions on large datasets using Apache Kylin technology.
from 2,950 PLNAzure Data Engineer Associate (DP-700)
This training prepares participants for the Microsoft Certified Azure Data Engineer Associate certification exam DP-700, which validates the skills needed to design and implement data engineering solutions on Microsoft Fabric. The course covers the full scope of the DP-700 exam objectives — including Lakehouse architecture, data pipeline development, semantic model creation, and real-time analytics — through a combination of conceptual instruction and extensive hands-on lab exercises. Participants will leave fully prepared to sit the DP-700 exam and to apply Microsoft Fabric data engineering skills in production environments.
from 4,500 PLNAzure for Data Engineers
Advanced Azure platform training focused on data processing and analysis. Participants will learn Azure tools and services for building data engineering solutions, from data integration to advanced analytics. Practical workshops cover implementation scenarios for big data and data lake solutions in Azure cloud.
from 3,750 PLNLarge-scale data analysis in Python
Training on large-scale data analysis in Python, covering pandas optimization, distributed processing with Dask and PySpark, and modern libraries Polars and Vaex. The program focuses on practical techniques for memory profiling, chunked processing, working with Parquet and Arrow formats, and database integration. Participants will learn strategies for choosing the right tools depending on data scale and characteristics.
from 2,450 PLNLarge-scale data analysis in R
Training on large-scale data analysis in R, covering memory optimization, parallel processing and integration with databases and Apache Spark. The program focuses on practical use of data.table, dplyr, future, foreach and sparklyr packages for efficient work with datasets that exceed the capabilities of standard tools. Participants will learn chunked processing techniques, memory profiling and query optimization.
from 2,450 PLNMicrosoft Fabric — Analytics Platform
This training provides a thorough introduction to Microsoft Fabric, Microsoft's unified analytics platform that brings together data engineering, data science, real-time analytics, and business intelligence in a single SaaS environment. Participants will learn to work with OneLake, Data Factory, Synapse Analytics, Power BI, and the AI-powered Copilot capabilities integrated across the platform. The course covers architecture, data pipelines, lakehouses, and end-to-end analytics workflows using practical hands-on labs.
from 3,500 PLNSMACK stack for Data Science
The training provides a practical introduction to the SMACK (Spark, Mesos, Akka, Cassandra, Kafka) technology stack in the context of Data Science solutions. The program combines theory with intensive workshops, during which participants learn how to integrate the various components and use them effectively in analytical projects. Classes are conducted in a workshop format with an emphasis on practical applications and real business scenarios.
from 2,950 PLNAI in data analysis and desk research
The training focuses on the use of artificial intelligence in data analysis and market research. Participants will learn to automate desk research, analyze trends and process large sets of information using AI. The program includes tools for Big Data mining, information filtering and forecasting market changes. An additional module is devoted to detecting fake news and misinformation, enabling effective assessment of source credibility and avoiding information manipulation.
from 2,450 PLNApache Hama - large-scale graph processing
Specialized training on graph processing in a distributed environment using Apache Hama. The program covers the foundations of graph theory, implementation of graph algorithms and techniques for optimizing computation in a Big Data environment. During the workshop, participants work with real graph datasets, learning to design and implement effective analytical solutions. The training combines mathematical theory with practical implementation aspects, focusing on the performance and scalability of solutions.
from 2,450 PLNApache SystemML in machine learning
The training provides a comprehensive introduction to Apache SystemML, a tool for big data processing and machine learning. During the workshop, participants will learn the principles of the system, its integration with big data platforms, and the implementation of machine learning algorithms. The program focuses on practical applications, using real-world use cases and large data sets.
from 2,450 PLNOperationalize Cloud Analytics Solutions with Microsoft Azure
552242A is a two-day instructor-led course designed for data professionals who want to expand their knowledge of creating big data analytics solutions in Microsoft Azure. Students will learn how to operationalize complex analytics solutions in the cloud using Azure Portal and Azure PowerShell. It can be used alone or with 552241A, Microsoft Azure Big Data Analytics, to prepare for Exam 70-475.
from 1,399 PLNSamza - Data Stream Processing
This training focuses on practical use of Apache Samza in real-time data stream processing. Through practical workshops, participants will learn the architecture and components of the Samza system and gain skills in designing and implementing stream processing solutions. Sessions are conducted in workshop format, with 70% of time dedicated to practical exercises.
from 2,450 PLNScientific Computing with SciPy Library
The training provides practical introduction to scientific computing in Python using the SciPy library. Participants will learn to effectively use available tools for solving complex mathematical problems and data analysis. The program is tailored to the needs of practitioners, combining theoretical foundations with intensive workshops solving real computational problems.
from 1,850 PLNZeppelin in interactive data analysis
A two-day advanced training dedicated to using Apache Zeppelin in interactive data analysis. The training program is designed to enable participants to effectively use Zeppelin notebooks in analytical projects, from data exploration to creating interactive dashboards. The training combines theory with intensive practical workshops where participants work on real use cases, learning best practices and data work techniques in a multilingual environment.
from 2,450 PLNAdvanced
Python and Spark for Big Data (PySpark)
Advanced training combining the capabilities of Python and Apache Spark in the context of Big Data processing. The workshop program covers both the theoretical foundations of distributed data processing and the practical aspects of implementing Big Data solutions. Participants will learn the tools and techniques necessary to efficiently process large data sets in a distributed environment, with a focus on performance optimization and scalability.
from 3,300 PLNApache Spark for developers - large-scale data processing
Advanced Apache Spark training focusing on the practical aspects of data processing in distributed environments. The program covers both fundamental concepts of distributed processing and advanced techniques for optimizing and implementing complex data flows. The workshop is conducted in the form of intensive hands-on classes, where participants work on real data sets, implementing a variety of analytical scenarios. Special emphasis is placed on understanding the internal mechanisms of Spark and the ability to use them effectively in production projects.
from 3,900 PLNApache Spark in the cloud - scalability and performance
The training focuses on the practical use of Apache Spark in cloud environments, with particular emphasis on scalability and performance aspects. The program covers advanced optimization techniques, monitoring and resource management in distributed computing systems. The class is conducted in the form of workshops with real-life scenarios where participants work on a cloud cluster.
from 3,900 PLNApache Spark MLlib - large-scale machine learning
The training provides an in-depth understanding of the Apache Spark MLlib library in the context of machine learning on large datasets. The program covers both the theoretical foundations of the algorithms and the practical aspects of implementing them in a distributed environment. Participants will learn advanced techniques for optimizing and scaling ML models and best practices for implementing production solutions.
from 6,350 PLNApache Spark Streaming using Scala
The training provides an in-depth understanding of Apache Spark Streaming using the Scala language. The program focuses on practical implementation of stream processing systems, with an emphasis on performance and reliability. Participants will learn advanced reactive programming techniques and design patterns used in real-time systems. The workshop includes working with real data streams.
from 3,300 PLNBig Data Hadoop analytics
The training prepares participants to work effectively as Big Data analysts in a Hadoop environment. The program combines technical knowledge with practical analytical skills, enabling participants to independently conduct advanced analysis on large data sets. Hands-on workshops focus on real-world use cases, allowing participants to gain experience working with the analytical tools of the Hadoop ecosystem. The classes are conducted in an interactive format with an emphasis on practical application of the knowledge gained.
from 4,500 PLNReal-Time Analytics - Apache Kafka, Flink and Stream Processing
Training on building advanced real-time analytics pipelines that deliver immediate business insights using enterprise-wide stream processing. The program includes practical application of Apache Flink, Kafka Streams and complex event processing techniques. Participants gain hands-on experience in implementing a modern streaming architecture. The stream-first methodology ensures a complete understanding of real-time data processing principles.
from 5,600 PLNAdvanced Apache Kafka — Streaming Data Architecture
Advanced Apache Kafka training covering event-driven architecture, Kafka Streams and ksqlDB, Kafka Connect (source/sink connectors), Schema Registry and Avro/Protobuf, exactly-once semantics, multi-datacenter replication (MirrorMaker 2), JMX and Grafana monitoring, and security (TLS, SASL, ACL).
from 4,500 PLNBig Data Integration with Talend
The training provides in-depth knowledge about data integration in Big Data environments using the Talend platform. The program is designed so participants can learn both basic and advanced tool features in the context of real integration challenges. Practical workshops allow gaining experience in designing, implementing, and managing data integration processes. The course pays particular attention to performance, scalability, and best practices in data flow design.
from 4,500 PLNBuilding Kafka solutions with Confluent
Advanced Confluent platform training that presents a comprehensive approach to building Apache Kafka-based solutions. The training program guides you through the process of designing and implementing data stream processing systems using the tools and components of the Confluent ecosystem. Hands-on workshops form the main part of the training, during which participants work on real implementation scenarios, learning about the platform's advanced capabilities and best practices for its use. The teaching methodology is based on a practical approach to problem solving, with an emphasis on production aspects.
from 2,450 PLNBusiness Intelligence with Big Data in Criminal Analysis
The training combines advanced Business Intelligence and Big Data techniques in the context of crime analysis. The program is designed to meet the specific requirements of law enforcement and security services. Participants learn through hands-on workshops to use modern analytical tools to detect criminal patterns and support investigations. The classes place special emphasis on data security, compliance with procedures and effective use of available information.
from 5,050 PLNCreating data pipelines with Apache Kafka
Intensive training focused on the design and implementation of data pipelines using Apache Kafka. The program covers advanced techniques for creating data flows, integrating sources and implementing real-time transformations. Participants learn through hands-on workshops to design efficient and scalable solutions for processing data streams. The training emphasizes real-world implementation scenarios and industry best practices.
from 1,850 PLNData Science in Big Data analysis
The training provides an advanced approach to Big Data analysis using Data Science methods. The program combines statistical theory with practical applications in a large-scale environment. Participants work on real data sets, learning to implement advanced analytical models. The class covers the full life cycle of a Data Science project, from data preparation to implementation of models in a production environment.
from 6,350 PLNDistributed messaging systems with Apache Kafka
Advanced training on the design and implementation of distributed messaging systems using Apache Kafka. The training program guides participants through the process of building scalable and reliable messaging systems, with a focus on performance and operational aspects. Hands-on workshops are a key component of the training, during which participants work with real implementation scenarios, learning best practices and design patterns. The teaching methodology is based on a gradual introduction of increasingly advanced concepts, allowing a full understanding of each system component.
from 2,450 PLNKafka for administrators - management and monitoring
Advanced training for Apache Kafka system administrators, focusing on the operational aspects of the platform. The program covers a comprehensive approach to managing, monitoring and maintaining Kafka clusters in a production environment. Hands-on workshops are the main component of the training, where participants work with real-world operational scenarios, learning to solve common problems and optimize system performance. The training places particular emphasis on the reliability and performance aspects of the platform.
from 3,900 PLNMicroservices with Spring Cloud and Kafka
Advanced training combining best practices for creating microservices in Spring Cloud with the capabilities of Apache Kafka. The program guides participants through the process of designing and implementing a scalable microservice architecture using event-driven patterns. During an intensive hands-on workshop, participants build a complex system, learning about integration patterns, asynchronous communication techniques and methods for ensuring fault tolerance. The training uses real-world design scenarios to present practical solutions to common challenges in microservice architectures.
from 3,300 PLNProcessing streams with Kafka Streams
Intensive training focused on the practical use of Kafka Streams for real-time processing of data streams. The program covers the design and implementation of processing pipelines, use of the Streams API, and stream processing patterns. The hands-on workshop allows participants to gain experience in developing high-performance streaming applications, taking into account aspects of scalability and reliability. The training uses real use cases to demonstrate practical applications of the technology.
from 1,850 PLNData Mesh Architecture - Domain Controlled Data and Self-Service Platforms.
The training presents a revolutionary approach to data management in large organizations, involving a shift from a centralized data lake model to decentralized data products focused on specific business domains. The program combines aspects of organizational transformation with the practical use of advanced technology tools. Participants will learn practical methods for implementing key principles of data mesh architecture. The domain-first methodology guarantees a sustainable and effective transformation of an organization's data architecture.
from 6,200 PLNDatameer for Data Analysts
Practical Datameer platform training designed for data analysts. The course focuses on using the tool for big data analysis, data preparation and visualization, and creating advanced analyses. Participants learn full data workflow in Datameer environment, from import to creating interactive reports.
from 2,450 PLNHortonworks Data Platform (HDP) for administrators
Advanced training in Hortonworks Data Platform administration, focusing on the practical aspects of managing the Hadoop ecosystem. The program covers configuration, maintenance and optimization of all key components of the HDP platform. Participants work in a production-like environment, learning to solve real-world operational problems. Classes are conducted in a workshop format, where theory is immediately translated into practical administrative scenarios. The training methodology is based on building a comprehensive understanding of the ecosystem through hands-on experience with its components.
from 3,900 PLNFrequently Asked Questions
Where to start learning Big Data?
Start with Apache Spark or Hadoop training — these are the foundations of the Big Data ecosystem. SQL knowledge and basic Python or Scala are helpful.
Do Big Data courses require programming experience?
Basic courses require SQL knowledge and one programming language (Python or Java). Advanced courses require database experience.
Which Big Data certifications are worth getting?
Top certifications include Cloudera CDP, Databricks Certified, AWS Data Analytics Specialty and Google Professional Data Engineer.
When should I use NoSQL instead of SQL?
NoSQL fits: horizontal scaling (MongoDB, Cassandra), unstructured data (document stores), relationship graphs (Neo4j), cache/sessions (Redis), real-time analytics (ClickHouse). Use SQL (PostgreSQL, MySQL) for: ACID transactions, reporting, structured business modeling. Modern stacks often combine both (polyglot persistence).
Can't find the right training?
We design custom training programs. Describe your needs and we'll prepare a tailored offer.
Ask about training