
Kick off your data science journey with a guided tour of the four course sections and course pathways, covering data mining, modeling, regression, and data preparation.
Explore why data science is the profession of the future as data grows exponentially, creating vast opportunities and a rising demand for skilled data scientists.
Choose from data-science pathways to tailor your learning, including independent modules, full monty, sequential, explorer, modeler, senior, and reporting services, covering visualization, data mining, modeling, data prep, communication and presentation.
Learn how to use ChatGPT to boost your Data Science skills and become more efficient!
Explore data visualization as a powerful tool for data mining and reporting using Tableau, with practical tutorials on AB testing and chi-squared tests in a drag-and-drop BI workflow.
Install Tableau Desktop or Tableau Public, choose the 14-day free trial or a free public version, download, install, and launch the desktop icon to begin analyzing.
Download office supplies data set from super data science training site, save it in visualization folder, and view 43-row CSV in Notepad++ or Excel to analyze regional sales in Tableau.
Learn to connect Tableau to a CSV data file, preview the data, choose live connections, and explore adding multiple data sources for later joins.
Navigate the Tableau workspace and distinguish dimensions from measures as independent and dependent variables. Build visuals by dragging data into rows and columns, using Show Me for quick chart ideas.
Learn how to use colors in Tableau to enhance data storytelling, including creating calculated fields, choosing color palettes, and mapping colors to regions for clear visual communication.
Learn how to export a Tableau worksheet as an image for Word or PowerPoint by copying the image, choosing view title and legend, and aligning sheet naming and formatting.
Connect Tablo to a CSFB file, navigate its intuitive menus, and distinguish measures from dimensions; then create a calculated field, add colors and labels, and export the worksheet for PowerPoint.
Investigate a 10,000-customer bank churn dataset to perform behavioral segmentation, assess predictors like gender and geography, and prepare for churn modeling with Tableau and AB testing basics.
Connect an Excel file to Tableau and visualize customer geography as a country map, using number of records to color by country and label counts.
Learn how to do an AB test in Tableau with accessible and comprehensive visualization
Learn to extend data mining from categorical to numeric variables, using visualization with age and balance, and apply the chi square test to back findings, with optional statistics tutorials.
Develop age-based bins to transform numeric data into categories, visualize customer distribution by age, and compare aggregated counts with percent-of-total to reveal a right-skewed pattern.
Create a classification test using a new age bin variable and visualize an age distribution chart with percentages, then plan to combine charts to assess significant results.
combine two Tableau charts to analyze age distribution and risk together, adding number of records, adjusting colors and labels, and comparing charts on a single worksheet to derive insights.
Explore the practical application of the chi-squared test to assess statistical significance in Tableau data mining, compare categories in AB tests, and validate insights with online tools.
Learn to apply chi-square tests to multi-category data, compare country groups, and validate findings with both a proper method and a quick approach.
Learn how the chi-squared test assesses independence by comparing observed and expected tables, calculates p-values, and applies Excel and online tools to verify results.
Apply the chi-squared test to assess independence between categorical variables using observed and expected tables. Convert data to absolute values, ensure mutually exclusive categories, and check p-values and cell counts.
Learn to create bins that convert numeric variables into categories, visualize distributions, and run AB tests for numeric data, using Tableau to combine charts and verify with chi-squared tests.
Finish part one of the epic data science course with visualization for data mining, and explore the next parts—modeling, data cleaning, or presentation—plus a Tableau bonus coupon.
Explore the modeling focus of data science, from intuition of logistic and linear regression to building and testing full statistical models, including geo demographic segmentation and post-delivery maintenance.
Explore the basic stats essential for the modeling part of the course, including types of parables, types of regressions, and R-squared and adjusted R-squared.
Explore the types of variables, including categorical versus numeric, with nominal and ordinal within categorical, and discrete and continuous within numeric, using practical examples.
Compare observed salaries with the regression line to understand how the model predicts outcomes. Minimize the sum of squared residuals with ordinary least squares to identify the best fitting line.
Learn how adjusted R-squared penalizes adding not-helpful variables in multiple regression, and compare it to R-squared to assess model robustness.
Explore Gretel, the free GUI for regression econometrics and time series modeling, with straightforward Windows and Mac installation and a no-code workflow.
Download and explore a salary dataset with two columns—years of experience and salary—use simple linear regression to model the relationship and demonstrate the business value of data-driven salary rules.
Import data in Gretel from a CSV, then view mean, median, min, max, and other summary statistics for variables, edit values, and save as Gretel data files.
Explore how to run an ordinary least squares regression, interpret the salary intercept and the years of experience coefficient, and assess p-values and R squared in the model report.
Plot and compare the actual versus fitted salaries using the linear regression graph and forecast tool; interpret 95% confidence intervals, the slope, and extrapolate salaries for new experience levels.
learn how to handle categorical variables with dummy variables, avoid the dummy variable trap, apply forward, backward, and bidirectional elimination to build regression models, and interpret results using adjusted R-squared.
Use multiple linear regression on a 50-company dataset to predict profit from r&d, administration, and marketing spend across states.
Examine the six linear regression assumptions, illustrated by the Anscombe quartet: linearity, equal variance, multivariate normality, independence, lack of multi collinearity, and an extra outlier check.
Use dummy variables to encode the state category (New York vs California) in a multiple linear regression predicting profit from R&D, admin, and marketing spend.
Explore the dummy variable trap in linear regression, highlighting multicollinearity when including multiple dummies. Learn to include only one dummy per set to ensure model clarity.
Understand how p values define statistical significance and drive hypothesis testing. See how rejecting the null hypothesis with a chosen confidence level uses a coin-toss example.
Explore step-by-step model building with backward elimination, forward selection, bidirectional elimination, all-in, and score comparison, using p-values and a 5% significance level to avoid garbage in, garbage out.
Learn backward elimination in Gretl to build a multiple linear regression with state dummies, 0.05 significance, and iterative predictor removal.
Explore how adjusted R-squared guides robust model selection beyond backward elimination, using penalization to decide whether to keep variables, and consider Akaike and Schwarz criteria.
This lecture recap covers dummy variables for categoricals, avoiding the dummy trap with n-1 dummies, backward, forward, bi-directional, and all-possible methods, and adjusted r-squared with unit-based coefficient interpretation.
Extremely Hands-On... Incredibly Practical... Unbelievably Real!
This is not one of those fluffy classes where everything works out just the way it should and your training is smooth sailing. This course throws you into the deep end.
In this course you WILL experience firsthand all of the PAIN a Data Scientist goes through on a daily basis. Corrupt data, anomalies, irregularities - you name it!
This course will give you a full overview of the Data Science journey. Upon completing this course you will know:
This course has pre-planned pathways. Using these pathways you can navigate the course and combine sections into YOUR OWN journey that will get you the skills that YOU need.
Or you can do the whole course and set yourself up for an incredible career in Data Science.
The choice is yours. Join the class and start learning today!
See you inside,
Sincerely,
Kirill Eremenko