2nd Edition

Introduction to Data Science Two-Volume Set

By Rafael A. Irizarry Copyright 2027
826 Pages 254 Color & 158 B/W Illustrations
by Chapman & Hall

826 Pages 254 Color & 158 B/W Illustrations
by Chapman & Hall

Unlike the first edition, the new edition has been split into two books, which have been brought together in this set. Thoroughly revised and updated, the first book ( Introduction to Data Science: Data Wrangling and Visualization with R ) introduces skills that can help the reader tackle real-world data analysis challenges. These include R programming, data wrangling with dplyr, data... Read more

Volume 1
Introduction Part 1: R 1. Getting started 2. R basics 3. Programming basics 4. The tidyverse 5. data.table 6. Importing data Part 2: Data Visualization 7. Visualizing data distributions 8. ggplot2 9. Data visualization principles 10. Data visualization in practice Part 3: Data Wrangling 11. Reshaping data 12. Joining tables 13. Parsing dates and times 14. Locales 15. Extracting data from the web 16. String processing 17. Text analysis Part 4: Productivity Tools 18. Organizing with Unix 19. Git and GitHub 20. Reproducible projects

Volume  2
Part 1: Summary Statistics 1. Distributions 2. Nummercial Summaries 3. Comparing Groups Part 2: Probability 4. Connecting Data and Probability 5. Discrete Probability 6. Continuous Probability 7. Random Variables 8. Sampling Models and the Central Limit Theorem Part 3: Statistical Inference 9. Sampling Models and the Central Limit Theorem 10. Data-Driven Models 11. Bayesian Statistics 12. Hierarchical Models 13. Hypothesis Testing 14. Bootstrap Part 4: Linear Models 15. Introduction to Regression 16. The Linear Model Framework 17. Treatment Effect Models 18. Generalized Linear Models 19. Association Is Not Causation 20. Multivariable Regression Part 5: High Dimensional Data 21. Working with Matrices in R 22. Applied Linear Algebra 23. Dimension Reduction 24. Regularization 25. Latent Factor Models Part 6: Machine Learning 26. Notation and Terminology 27. Performance Metrics 28. Conditional Expectations and Smoothing 29. Resampling and Model Assessment 30. Supervised Learning Methods 31. Building Machine Learning Models 32. Unsupervised Learning: Clustering

Biography

Rafael A. Irizarry is Professor and Chair of the Department of Data Science at Dana-Farber Cancer Institute and Professor of Applied Statistics at Harvard. His research focuses on Genomics and he has taught several Data Science courses.