Skip to content
Technologies

RAG and LangChain — Building AI Applications with Custom Data

A hands-on training on building AI applications powered by Retrieval-Augmented Generation (RAG) and the LangChain framework. The program covers RAG architectures, chunking and embedding strategies, vector databases (Pinecone, ChromaDB, FAISS), advanced retrieval techniques including hybrid search and reranking, and building production-ready RAG pipelines. Participants will learn to evaluate response quality using RAGAS and deploy a complete RAG system from prototype to production.

Why choose this training?

Retrieval-Augmented Generation is currently the most effective approach for building AI applications that answer questions based on organizational data — without hallucinations and without expensive fine-tuning. However, building a RAG system that works reliably in production requires far more than connecting documents to a vector database. Critical decisions around chunking strategy, embedding model selection, retrieval configuration, and response quality evaluation determine the difference between a demo and a production system.

This two-day training guides participants through the entire RAG development lifecycle — from document splitting and embedding creation, through vector database configuration, to advanced hybrid search and reranking techniques. Every module includes hands-on exercises in a prepared lab environment, where participants independently build and test individual RAG pipeline components using LangChain.

What makes our approach unique?

EITT tracks the rapidly evolving RAG ecosystem and frameworks like LangChain, regularly updating the curriculum with the latest retrieval techniques, embedding models, and evaluation tools. Our instructors have hands-on experience deploying production RAG systems in organizations of varying scale — from startups to enterprise. With a team of over 500 experts and experience from more than 2,500 trainings, we deliver programs grounded in proven patterns rather than experiments.

The training places particular emphasis on production readiness — participants not only build a prototype but learn to scale, monitor, and optimize it. The program covers RAGAS evaluation, token cost management, and knowledge base update strategies that are critical in real-world enterprise deployments.

Benefits

  • Understand RAG architecture and its advantages over LLM fine-tuning
  • Build a complete RAG pipeline with ingestion, retrieval, and generation
  • Select and configure the right vector database for the project scale
  • Implement advanced retrieval techniques with hybrid search and reranking
  • Build a production RAG system with caching, guardrails, and monitoring
  • Evaluate RAG system quality using RAGAS and other evaluation metrics
  • Integrate RAG with LangChain to handle multiple data sources

Who is this training for?

Python developers building AI-powered applications with custom data
ML engineers seeking practical RAG implementation patterns
Data scientists integrating LLMs with organizational knowledge bases
AI architects designing semantic search and retrieval systems
Backend developers extending applications with AI capabilities
Technical leads evaluating RAG adoption for enterprise use cases

Prerequisites

  • Intermediate Python proficiency (minimum 1 year of experience)
  • Basic understanding of large language models and APIs (OpenAI, Anthropic, or similar)
  • Experience with REST APIs and JSON
  • Basic understanding of NLP and ML concepts (embeddings, similarity search)

Training program

01

Module 1: RAG Fundamentals — Architecture, Chunking, and Embedding

  • What is RAG and why it solves the LLM hallucination problem
  • RAG pipeline architecture — ingestion, retrieval, and generation
  • Chunking strategies — fixed-size, semantic, recursive, and document-aware
  • Embedding models — OpenAI, Cohere, open-source alternatives and comparison
  • Impact of chunking and embedding quality on answer accuracy
  • Hands-on — building a first RAG pipeline from scratch
02

Module 2: LangChain Framework — Chains, Agents, and Memory

  • LangChain architecture — components, chains, and orchestration
  • Prompt templates and output parsers in the RAG context
  • Chains vs agents — choosing the right approach
  • Memory systems in LangChain — conversation buffer, summary, and vector memory
  • Integrating LangChain with external data sources
  • Debugging and tracing with LangSmith
03

Module 3: Vector Databases in RAG — Pinecone, ChromaDB, and FAISS

  • Vector database landscape — managed vs self-hosted options
  • Pinecone — configuration, indexing, and metadata filtering
  • ChromaDB — local vector database for prototyping and production
  • FAISS — high-performance vector search at scale
  • Indexing strategies and data partitioning approaches
  • Performance comparison and selection criteria for vector databases
04

Module 4: Advanced Retrieval — Hybrid Search, Reranking, and Metadata Filtering

  • Dense retrieval vs sparse retrieval — BM25, TF-IDF, and embeddings
  • Hybrid search — combining dense and sparse retrieval results
  • Reranking — Cohere Rerank, cross-encoders, and ColBERT
  • Metadata filtering — narrowing results by document attributes
  • Query transformation — HyDE, multi-query, and step-back prompting
  • Practical scenarios for choosing the right retrieval strategy
05

Module 5: Building Production RAG Pipelines

  • Production architecture — scaling, caching, and load balancing
  • Knowledge base updates — incremental indexing and versioning
  • Handling multiple data sources — PDFs, web pages, SQL databases, APIs
  • Guardrails — input and output validation in the RAG pipeline
  • Cost management — token optimization and caching strategies
  • RAG integration patterns for existing enterprise systems
06

Module 6: Evaluation and Optimization — RAGAS, Faithfulness, and Relevance

  • RAGAS framework — evaluation metrics for RAG systems
  • Faithfulness — verifying answers stay grounded in context
  • Answer relevance and context relevance — retrieval accuracy
  • Building test datasets and benchmarks
  • Iterative pipeline optimization based on evaluation metrics
  • Production monitoring — quality alerts and RAG dashboards

Delivery Methods

Online

  • Convenience of participating from anywhere
  • Interactive live sessions with trainer
  • Materials available for 30 days
  • No travel costs

On-site

  • Direct contact with trainer and group
  • Intensive hands-on workshops
  • Networking with other participants
  • Full focus on learning

Frequently asked questions

How does RAG differ from fine-tuning, and when should I choose RAG?

RAG allows models to access up-to-date, organization-specific data without retraining. Fine-tuning permanently alters model behavior and requires an expensive training process, while RAG dynamically provides context from a knowledge base at query time. RAG is the better choice when data changes frequently, when source transparency is required, or when the budget does not allow for fine-tuning.

Which vector databases are covered and how do I choose the right one?

The training covers three vector databases: Pinecone (managed, best for production without infrastructure management), ChromaDB (open-source, ideal for prototypes and smaller deployments), and FAISS (Facebook's library, most performant for large-scale on-premise datasets). The choice depends on project scale, latency requirements, budget, and hosting preferences — these criteria are discussed in detail during the hands-on module.

What types of data can I connect to a RAG system?

A RAG system can process virtually any text-based data — PDF documents, web pages, SQL databases, Markdown files, wiki articles, API documentation, and email messages. The training demonstrates specific loaders and chunking strategies for different source types, including multilingual documents and structured data.

Will I be able to deploy RAG in production after this training?

Yes — the training deliberately covers the full production lifecycle, not just prototyping. Production modules address scaling, caching, incremental knowledge base updates, guardrails, response quality monitoring, and token cost management. Participants finish the training with a working RAG pipeline and the knowledge needed to deploy it in their organization.

Why choose EITT for this RAG and LangChain training?

EITT is a training provider with over 500 experts and experience from more than 2,500 delivered trainings. Our RAG and LangChain instructors have hands-on experience deploying production AI systems in enterprise organizations. The curriculum is regularly updated to reflect the latest framework versions and vector database capabilities. EITT maintains an average participant rating of 4.8/5.

Adrian Kwiatkowski
Adrian Kwiatkowski Opiekun szkolenia

Request a quote

Funding Options

Check funding options for your company

Up to 80%

Development Services Database

Up to 80% funding for SMEs from EU funds

Check availability
Up to 100%

National Training Fund

Up to 100% funding for employers

Learn more

Trusted by

We train teams at Poland's largest companies

ING Bank - EITT client
mBank - EITT client
PKO Bank Polski - EITT client
PZU - EITT client
Allianz - EITT client
T-Mobile - EITT client
KGHM - EITT client
PGE - EITT client
IKEA - EITT client
InPost - EITT client
Leroy Merlin - EITT client
ZUS - EITT client

Interested in this training?

Contact us - we'll prepare an offer tailored to your organization's needs.

500+ experts
2500+ trainings available
ISO 9001 quality certified
Request Training
Call us +48 22 487 84 90