Các bạn đăng ký học viên để truy cập file này có kèm hyperlink liên kết trực tiếp với tài liệu gốc giúp thuận tiện tra cứu thông tin.

Liên hệ: tuhocr.com@gmail.com

Cơ sở lý thuyết về trực quan hóa dữ liệu trong R

Sơ đồ tương quan các hệ thống plotting trong R

Example base-R graphic

Đặc điểm của các lệnh trong base-R là lệnh plot() sẽ tạo đồ thị, sau đó các lệnh phụ như abline() hay text() sẽ chèn thêm thông tin cần thiết, theo thứ tự từ trên xuống.

fit_lm <- lm(mpg ~ wt, data = mtcars)
plot(mpg ~ wt, data = mtcars, pch = 19, col = 4, 
     ylim = c(0, 40), xlim = c(0, 10))
abline(fit_lm, col = 2, lwd = 2)
text(x = 6, y = max(mtcars$mpg)-4,  
     bquote(atop(Y == ~ .(round(fit_lm$coefficients[1], 4)) + 
                     (.(round(fit_lm$coefficients[2], 4))) * "\u00D7" * X, 
                 R^2 == .(round(unclass(summary(fit_lm))$r.squared, 4)))),
     cex = 1.5)

Example lattice package

Đặc điểm của các lệnh trong lattice là toàn bộ format sẽ nằm trong lệnh plot tương ứng, không tách ra từng lệnh phụ riêng lẻ như base-R graphic.

library("lattice")
dotplot(variety ~ yield | site, data = barley, groups = year,
        key = simpleKey(levels(barley$year), space = "right"),
        xlab = "Barley Yield (bushels/acre) ",
        aspect = 0.5, layout = c(1, 6), ylab=NULL,
        scales = list(tck = c(1, 0), x = list(cex = 0.8), y = list(cex = 0.65)),
        par.strip.text = list(cex=0.9)
        )

Example ggplot2 package

Đặc điểm các lệnh trong ggplot2 là đưa thông tin vào đồ thị theo kiểu layer/lớp, mỗi lớp được đưa vào sử dụng dấu +, nếu cần xuống hàng thì xuống hàng ngay sau dấu + để R biết được là các lệnh này thuộc cùng 1 câu lệnh chung.

options(width = 200)
## Edgar Anderson's Iris Data
## Description
# This famous (Fisher's or Anderson's) iris data set gives 
# the measurements in centimeters of the variables sepal length and width and petal length and width, 
# respectively, for 50 flowers from each of 3 species of iris. 
# The species are Iris setosa, versicolor, and virginica.

head(iris)
##   Sepal.Length Sepal.Width Petal.Length Petal.Width Species
## 1          5.1         3.5          1.4         0.2  setosa
## 2          4.9         3.0          1.4         0.2  setosa
## 3          4.7         3.2          1.3         0.2  setosa
## 4          4.6         3.1          1.5         0.2  setosa
## 5          5.0         3.6          1.4         0.2  setosa
## 6          5.4         3.9          1.7         0.4  setosa
str(iris)
## 'data.frame':    150 obs. of  5 variables:
##  $ Sepal.Length: num  5.1 4.9 4.7 4.6 5 5.4 4.6 5 4.4 4.9 ...
##  $ Sepal.Width : num  3.5 3 3.2 3.1 3.6 3.9 3.4 3.4 2.9 3.1 ...
##  $ Petal.Length: num  1.4 1.4 1.3 1.5 1.4 1.7 1.4 1.5 1.4 1.5 ...
##  $ Petal.Width : num  0.2 0.2 0.2 0.2 0.2 0.4 0.3 0.2 0.2 0.1 ...
##  $ Species     : Factor w/ 3 levels "setosa","versicolor",..: 1 1 1 1 1 1 1 1 1 1 ...
summary(iris)
##   Sepal.Length    Sepal.Width     Petal.Length    Petal.Width          Species  
##  Min.   :4.300   Min.   :2.000   Min.   :1.000   Min.   :0.100   setosa    :50  
##  1st Qu.:5.100   1st Qu.:2.800   1st Qu.:1.600   1st Qu.:0.300   versicolor:50  
##  Median :5.800   Median :3.000   Median :4.350   Median :1.300   virginica :50  
##  Mean   :5.843   Mean   :3.057   Mean   :3.758   Mean   :1.199                  
##  3rd Qu.:6.400   3rd Qu.:3.300   3rd Qu.:5.100   3rd Qu.:1.800                  
##  Max.   :7.900   Max.   :4.400   Max.   :6.900   Max.   :2.500

Vẽ đồ thị scatter plot

library("ggplot2")
library("dplyr")

ggplot(iris, aes(x = Sepal.Width, y = Sepal.Length, color = Species)) + 
    geom_point(size=2) +
    labs(col = "Species") +
    scale_color_manual(labels = c("setosa", "versicolor", "virginica"),
                       values = c("red", "darkgreen", "blue")) +
    theme_classic()

Vẽ đồ thị scatter plot phân nhóm theo từng loài

iris.sc <- iris
iris.sc[,1:4] <- scale(iris[,1:4])
iris.sc.input <- iris.sc[,1:4]

set.seed(123) # Set seed number
irisKM.mod <- kmeans(iris.sc.input, centers = 3, nstart = 100)

groupPred <- factor(irisKM.mod$cluster, levels = c(1,2,3), ordered = FALSE)
iris$KMpred <- groupPred

groupPred <- factor(irisKM.mod$cluster, levels = c(1,2,3), ordered = FALSE)
iris$KMpred <- groupPred

# Plot the Data
ggplot(iris, aes(y = Sepal.Length, x = Sepal.Width, col = KMpred)) +
    geom_point(size=2) +
    stat_ellipse(level = 0.9) +
    labs(col = "Species") +
    scale_color_manual(labels = c("setosa", "versicolor", "virginica"),
                       values = c("red", "darkgreen", "blue")) +
    theme_classic()

Tài liệu tham khảo

Jones, E., Harden, S., & Crawley, M. J. (2023). The R Book (3rd ed.). Wiley.
Murrell, P. (2019). R Graphics (3rd ed.). CRC Press, Taylor & Francis Group.
Sarkar, D. (2008). Lattice: Multivariate Data Visualization with R. Springer New York.

Tham gia khóa học

Để học R bài bản từ A đến Z, thân mời Bạn tham gia khóa học “HDSD R để xử lý dữ liệu”

ĐĂNG KÝ NGAY: https://www.tuhocr.com/register