Data science bootcamp online
A comprehensive program covering data science fundamentals, visualization, predictive modeling, model deployment, and advanced topics like text analytics and recommender systems.
Trusted by leading companies




Who is this bootcamp for?
Aspiring data scientists & analysts
Learn the full stack: Python, data processing, statistics, and machine learning. Build real projects, not just theory.
Engineers & technical professionals
Expand your toolkit and learn to apply statistical models and machine learning to solve complex engineering and business problems.
Career changers
Whether you come from finance, operations, IT, or another analytical field, this program provides the structure and skills to transition into data roles confidently.
Meet the instructors
Learn from practitioners who build and deploy AI systems at scale
Curriculum
16 modules covering the full spectrum of theory and practice
Getting ready for the bootcamp
In this module, we'll build the foundational programming knowledge and theory needed to succeed in the bootcamp. By covering essential Python concepts and tools, we'll set ourselves up for a smoother learning experience during the hands-on sessions. Whether you're starting from scratch or refreshing your skills, this module is the perfect starting point for mastering Python and its practical applications.
What you'll learn
- Grasp the fundamentals of Python programming, including key concepts for beginners
- Master efficient input/output operations to handle data seamlessly
- Navigate and utilize Jupyter Notebook for writing and executing Python code
- Learn how to interact with REST APIs effectively using Python
Exploratory data analysis
In this module, we'll focus on data exploration, visualization, and feature engineering—essential steps in preparing data for analysis. We'll learn how to use techniques like summary statistics and visual tools to understand the structure of your data, spot patterns, and identify issues such as missing values or outliers. We'll also cover how to transform and create new features that make your data more useful for modeling. By the end, you'll be able to uncover insights, clean your data, and shape it for better analysis and predictive performance.
What you'll learn
- Explore the idea behind exploratory data analysis using Python
- Understand the importance of feature engineering in building effective machine-learning models
- Apply summary statistics to interpret data
- Analyze the importance of segmentation in data analysis and recognize pitfalls of data analysis
- Understand how domain knowledge helps us to create informative and relevant features
Storytelling with data
In this module, we'll explore how to turn raw data into compelling visual stories. We'll learn how to choose the right visualizations for different data types, and apply tools like pandas, matplotlib, and seaborn to analyze real datasets. Through hands-on exercises, we'll uncover key insights and develop a deeper understanding of the data. We'll also discuss how to design and present visuals that clearly communicate the message to our audience. By the end of the module, you'll be equipped to present data in a clear, impactful, and audience-friendly way.
What you'll learn
- Understand how to select appropriate visualization types for different types of data
- Analyze data using Pandas, Matplotlib, and Seaborn through hands-on exercises
- Unlock insights and gain a deeper understanding of data
- Discuss how to deliver effective visuals for data storytelling
Predictive modeling for real world problems
In this module, we'll explore how to use predictive modelling to create real business impacts. We'll learn to identify the right opportunities for machine learning, translate business goals into actionable models, and recognize when good models might still lead to poor outcomes. We'll also cover essential data requirements and key ethical considerations. By the end, you'll be able to understand predictive systems that are both effective and aligned with business strategy.
What you'll learn
- Learn how to identify the right business opportunities for real analytics impact
- Investigate conditions where a good machine learning model leads to adverse business impact
- Understand various types of data necessary for data collection and analysis
- Identify how to translate business goals into actionable machine learning solutions
- Understand key ethical and practical challenges in building effective predictive systems
Decision tree learning
In this module, we will explore decision tree learning and focus on how decision trees are constructed for supervised classification tasks. We will learn to apply splitting criteria like Gini index and how to evaluate model performance. Practical exercises and quizzes at the end will help you to apply the learned skills to real datasets and benchmark your performance. By the end of this module, you will be equipped to confidently implement decision tree classifiers.
What you'll learn
- Understand the fundamentals of decision tree learning
- Learn how to choose optimal split points using criteria like Gini index
- Explore key decision tree concepts such as depth, overfitting, and pruning
- Apply decision tree algorithms to real-world datasets through hands-on exercises
- Gain experience with model interpretation and visualizing tree-based decisions
Evaluation of classification models
In this module, we'll focus on evaluating classification models and understanding their impact in real-world applications. We'll begin by discussing why accuracy alone isn't a reliable metric and how different types of errors affect business decisions. Through the confusion matrix and key performance metrics like precision, recall, and F1 score, we'll learn to interpret model results with clarity. We'll also explore advanced evaluation tools like ROC curves and AUC to help us assess model performance across thresholds. By the end of this module, you'll be able to choose appropriate metrics, interpret model results effectively, and make evaluation decisions that align with real-world goals.
What you'll learn
- Explain why accuracy alone may be misleading when evaluating classification models
- Interpret the confusion matrix to assess classification performance
- Differentiate between key metrics such as precision, recall, and F1 score
- Evaluate metric trade-offs in different business contexts
- Apply ROC curves and AUC to compare model performance beyond threshold-based metrics
Tuning of model hyperparameters
Master the essentials of hyperparameter tuning to improve your model's accuracy and reliability. This module covers key concepts like generalization, overfitting, and the bias-variance tradeoff, alongside practical validation techniques and strategies to handle real-world data challenges. We'll learn how to fine-tune models for better performance and robust predictions.
What you'll learn
- Describe the role of hyperparameter tuning in improving machine learning model performance
- Identify and apply key techniques for effective hyperparameter tuning
- Explain what overfitting is and how it affects model generalization
- Recognize the characteristics of a well-generalized model across different data samples
- Understand the bias-variance tradeoff and how it influences tuning decisions
Ensemble methods, bagging and random forest
Unlock the potential of ensemble methods and elevate your predictive power. This module aims to go deep into techniques such as bagging and random forests, explore key concepts like the bias-variance trade-off and out-of-bag evaluation, and build practical skills through interactive notebooks. Whether you're working with messy real-world data or aiming for stronger model performance, ensemble learning is the next step forward!
What you'll learn
- Explore key techniques like bagging and random forests in depth
- Apply concepts such as out-of-bag evaluation and binomial probability
- Understand how ensemble methods balance bias and variance
- Gain hands-on experience through interactive, real-world exercises
Boosting
Boosting is a powerful ensemble technique that builds better models by learning from previous mistakes, one step at a time. In this course, we'll explore the core ideas behind boosting, explore popular methods like AdaBoost, and apply them through real-world examples and hands-on exercises. Whether you're just starting out or looking to sharpen your skills, this course will help you create smarter and more accurate predictive models.
What you'll learn
- Understand the concept of using weak classifiers to create a strong one
- Analyze boosting and its ability to adjust the weight of each classifier
- Apply the logic behind the boosting algorithm to real-world examples
- Evaluate the differences between bagging and boosting
- Discuss the potential drawbacks of the boosting algorithm
Online experimentation and A/B testing
In today's data-driven world, successful digital products are not built on guesswork. They're built on evidence. This module will walk you through the essential principles and practical techniques of online experimentation. From understanding the basics of A/B testing to mastering advanced testing strategies and avoiding common pitfalls, you'll gain the knowledge and confidence to run meaningful experiments that drive better outcomes. Whether you're refining a feature, testing a new idea, or scaling insights across teams, this module sets the foundation for thoughtful, measurable, and impactful decisions.
What you'll learn
- Recognize the role of experimentation in online business and marketing
- Learn how to design and run effective experiments
- Distinguish between A/B, A/A, and multivariate testing and their use cases
- Apply best practices for metrics selection, hypothesis creation, and error avoidance
- Interpret results using statistical significance and confidence intervals
Text analytics fundamentals
Every tweet, email, review, and support ticket holds valuable insights if we know how to extract them. This module introduces the tools and techniques businesses use to turn raw text into strategic decisions. We will learn how to clean, structure, and analyze text using NLP, TF-IDF, and word embeddings like Word2Vec. Whether you want to understand customer sentiment or automate tasks, this module will help you get started.
What you'll learn
- Analyze unstructured data types such as text
- Evaluate the process of converting text into structured data
- Understand the steps involved in pre-processing text
- Apply the concepts of TF-IDF and Word2Vec
- Use cosine similarity to determine the similarity between documents
Unsupervised learning with K-means clustering
Discover how unsupervised learning helps uncover hidden patterns in data without labels. This module introduces key concepts of clustering, with a focus on the k-means algorithm. We'll learn how k-means works, when to use it, and how to choose the right number of clusters for meaningful insights from raw data.
What you'll learn
- Discuss the basics of unsupervised learning and how it differs from supervised learning
- Explain the concept and steps of the K-means clustering algorithm
- Apply K-means to group unlabeled data and uncover patterns
- Evaluate the strengths and limitations of K-means clustering
- Determine the optimal number of clusters using the elbow method
Linear models for regression
Discover how a simple straight line can unlock powerful insights and help predict the future. In this course, we'll explore how linear models work, how to evaluate their performance, and how optimization techniques like gradient descent help improve accuracy. Through intuitive explanations and practical examples, we'll learn to build, interpret, and apply linear regression models with confidence.
What you'll learn
- How linear regression models predict continuous outcomes using input features
- The difference between parametric and non-parametric approaches
- The role of cost functions in quantifying prediction error
- How gradient descent optimizes model parameters step by step
- How to measure model performance using different metrics
Regularization and tuning of linear models
What if our linear models could stay accurate and simple, even in messy real-world data? In this module, we will build models that balance precision and simplicity. We will learn how to control model complexity, use regularization to prevent overfitting, and tune hyperparameters so our models make reliable, real-world predictions.
What you'll learn
- Balancing model complexity to avoid underfitting and overfitting
- Applying L1 (Lasso) and L2 (Ridge) regularization to control complexity
- Tuning key hyperparameters for optimal performance
- Evaluating model fit and use cross-validation to improve generalization
- Interpreting the bias-variance trade-off when tuning linear models
Ranking and recommendation systems
Ever wondered how Netflix knows what we'll love next? This module unpacks the magic behind recommendation systems, from how they measure similarity and generate suggestions to choosing the right approach for your business. Discover how to use data-driven recommendations to boost engagement, delight customers, and drive results.
What you'll learn
- Different type of recommendation systems like collaborative and content-based
- Similarity measures to compare users or items and find the best matches
- Recommendation algorithms used to generate and rank suggestions
- Evaluation techniques to assess performance and effectiveness
Big data engineering with distributed systems
Learn the core principles of big data engineering and distributed systems to confidently tackle large-scale data challenges. This module introduces key topics such as cloud infrastructure models, distributed computing frameworks, and scalable system design. We'll use tools like Hadoop and Spark to process data efficiently and explore modern architectures that support real-time analytics and machine learning workflows.
What you'll learn
- Explain the role of data engineering in enabling scalable machine learning workflows
- Describe how distributed systems process and manage large-scale data efficiently
- Compare cloud service models and assess their suitability for different applications
- Identify the core components and workflows of Hadoop, Spark, and MapReduce
Earn a verified certificate
Earn a verified certificate and take your next career step with credibility.
Recognized by



Reserve your spot
We have carefully designed our data science bootcamp to bring you the best practical exposure in the world of data science, programming, and machine learning.
- Data science bootcampOnline
Self paced
Launching soon
What our alumni say
"The data science bootcamp is an amazing way to improve your skills in data science and engineering! The content of the training and the trainers themselves are amazing. Now after the bootcamp I am very inspired to go back to my data and implement all the new knowledge I gained! The bootcamp of Data Science Dojo is one I will recommend to all who are interested in becoming skilled in the amazing field of data science! - Amanda ten Brink attended Data Science and Data Engineering Bootcamp."
Frequently asked questions
- Are classes live or self-paced?
- Currently, this bootcamp is discontinued, as we plan to relaunch a self-paced format later this year. Once the updated structure and details are finalized, they will be shared here on our website.
- Is the program full-time or part-time, and in-person or online?
- The program is currently not offered in-person, and is self-paced online. To complete the program within five days, it usually requires a full-time commitment of 40 hours in total (8 hours per day). However, the self-paced program allows you to work on your own time, so you can complete it over several days in part-time hours that suit you.
- What is the cost and are discounts available?
- More information on pricing and discounts is yet to come, as we plan to relaunch a self-paced format of the bootcamp later this year. Once the updated structure and details are finalized, they will be shared here on our website.
- Is the training conducted in R or Python?
- We are a technology-neutral and vendor-agnostic training. Both R and Python code samples will be shared with attendees.
- Are recorded lesson clips available?
- All key topics and content are available in our companion courses. We've created structured lesson clips that reflect the material discussed, allowing you to review everything at your convenience.
- Can I reach out if I have questions while working on the course and homework?
- To support you, we have a dedicated Discord community where you can receive help from our instructors and connect with fellow students.
- Will I earn a certificate from this data science bootcamp?
- Yes! Once you have completed the Data science bootcamp, you will be issued a verified certificate of completion. You can print or add this certificate to your LinkedIn profile for others to see.
- What backgrounds do people have that take this data science bootcamp?
- We have had attendees from a wide range of backgrounds – software engineers, product/program managers, physicists, financial analysts – even medical doctors and veterinarians, attend and successfully completed our data science bootcamp. This bootcamp is for anyone who is curious about data science and willing to explore, segment, analyze, and understand their data in order to make better data-driven decisions.
- What kind of jobs can a data science bootcamp get me?
- Data Scientist, Data Engineer, Machine Learning Engineer, Data Analyst, Business Analyst / Product Analyst
- What is the transfer policy?
- Transfers are allowed once with no penalty. Transfers requested more than once will incur a $200 processing fee.
- What is the refund policy?
- If, for any reason, you decide to cancel, we will gladly refund your registration fee in full if you notify us at least five business days before the start of the training. We can also transfer your registration to another cohort if preferred. However, refunds cannot be processed if you have transferred to a different cohort after registration. Additionally, once you have been added to the learning platform and have accessed the course materials, we are unable to issue a refund, as digital content access is considered program participation.

