8/29/25, 4:31 PM Exploratory Data Aanalysis
Exploratory Data Aanalysis
Shahrizal MA
Exploratory Data Analysis
Langkah-langkah dalam mengeksplor dan memeriksa data sebelum membangun sebuah model dengan cara
sederhana, yaitu melihat struktur data, menghitung statistika 5 serangkai (statistik deskriptif), dan visualisasi data.
Berikut langkah-langkah yang diberlakukan pada data mtcars dan mpg .
Struktur Data
data("mtcars")
library(ggplot2)
data("mpg")
dd <- mtcars
df <- mpg
str(dd)
## '[Link]': 32 obs. of 11 variables:
## $ mpg : num 21 21 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 ...
## $ cyl : num 6 6 4 6 8 6 8 4 4 6 ...
## $ disp: num 160 160 108 258 360 ...
## $ hp : num 110 110 93 110 175 105 245 62 95 123 ...
## $ drat: num 3.9 3.9 3.85 3.08 3.15 2.76 3.21 3.69 3.92 3.92 ...
## $ wt : num 2.62 2.88 2.32 3.21 3.44 ...
## $ qsec: num 16.5 17 18.6 19.4 17 ...
## $ vs : num 0 0 1 1 0 1 0 1 1 1 ...
## $ am : num 1 1 1 0 0 0 0 0 0 0 ...
## $ gear: num 4 4 4 3 3 3 3 4 4 4 ...
## $ carb: num 4 4 1 1 2 1 4 2 2 4 ...
str(df)
## tibble [234 × 11] (S3: tbl_df/tbl/[Link])
## $ manufacturer: chr [1:234] "audi" "audi" "audi" "audi" ...
## $ model : chr [1:234] "a4" "a4" "a4" "a4" ...
## $ displ : num [1:234] 1.8 1.8 2 2 2.8 2.8 3.1 1.8 1.8 2 ...
## $ year : int [1:234] 1999 1999 2008 2008 1999 1999 2008 1999 1999 2008 ...
## $ cyl : int [1:234] 4 4 4 4 6 6 6 4 4 4 ...
## $ trans : chr [1:234] "auto(l5)" "manual(m5)" "manual(m6)" "auto(av)" ...
## $ drv : chr [1:234] "f" "f" "f" "f" ...
## $ cty : int [1:234] 18 21 20 21 16 18 18 18 16 20 ...
## $ hwy : int [1:234] 29 29 31 30 26 26 27 26 25 28 ...
## $ fl : chr [1:234] "p" "p" "p" "p" ...
## $ class : chr [1:234] "compact" "compact" "compact" "compact" ...
[Link] R publish/R pubs/[Link] 1/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
Statistika Deskriptif
summary(dd)
## mpg cyl disp hp
## Min. :10.40 Min. :4.000 Min. : 71.1 Min. : 52.0
## 1st Qu.:15.43 1st Qu.:4.000 1st Qu.:120.8 1st Qu.: 96.5
## Median :19.20 Median :6.000 Median :196.3 Median :123.0
## Mean :20.09 Mean :6.188 Mean :230.7 Mean :146.7
## 3rd Qu.:22.80 3rd Qu.:8.000 3rd Qu.:326.0 3rd Qu.:180.0
## Max. :33.90 Max. :8.000 Max. :472.0 Max. :335.0
## drat wt qsec vs
## Min. :2.760 Min. :1.513 Min. :14.50 Min. :0.0000
## 1st Qu.:3.080 1st Qu.:2.581 1st Qu.:16.89 1st Qu.:0.0000
## Median :3.695 Median :3.325 Median :17.71 Median :0.0000
## Mean :3.597 Mean :3.217 Mean :17.85 Mean :0.4375
## 3rd Qu.:3.920 3rd Qu.:3.610 3rd Qu.:18.90 3rd Qu.:1.0000
## Max. :4.930 Max. :5.424 Max. :22.90 Max. :1.0000
## am gear carb
## Min. :0.0000 Min. :3.000 Min. :1.000
## 1st Qu.:0.0000 1st Qu.:3.000 1st Qu.:2.000
## Median :0.0000 Median :4.000 Median :2.000
## Mean :0.4062 Mean :3.688 Mean :2.812
## 3rd Qu.:1.0000 3rd Qu.:4.000 3rd Qu.:4.000
## Max. :1.0000 Max. :5.000 Max. :8.000
summary(df)
## manufacturer model displ year
## Length:234 Length:234 Min. :1.600 Min. :1999
## Class :character Class :character 1st Qu.:2.400 1st Qu.:1999
## Mode :character Mode :character Median :3.300 Median :2004
## Mean :3.472 Mean :2004
## 3rd Qu.:4.600 3rd Qu.:2008
## Max. :7.000 Max. :2008
## cyl trans drv cty
## Min. :4.000 Length:234 Length:234 Min. : 9.00
## 1st Qu.:4.000 Class :character Class :character 1st Qu.:14.00
## Median :6.000 Mode :character Mode :character Median :17.00
## Mean :5.889 Mean :16.86
## 3rd Qu.:8.000 3rd Qu.:19.00
## Max. :8.000 Max. :35.00
## hwy fl class
## Min. :12.00 Length:234 Length:234
## 1st Qu.:18.00 Class :character Class :character
## Median :24.00 Mode :character Mode :character
## Mean :23.44
## 3rd Qu.:27.00
## Max. :44.00
[Link] R publish/R pubs/[Link] 2/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
Visualisasi Data
1. Density Plot
library(ggplot2)
library(hrbrthemes)
## Warning: package 'hrbrthemes' was built under R version 4.4.2
# data Transmisi dan Engine
data <- [Link](
var1 = dd$vs,
var2 = dd$am)
p1 <- ggplot(data, aes(x=x) ) +
# Top
geom_density( aes(x = var1, y = ..density..), fill="#69b3a2" ) +
geom_label( aes(x=4.5, y=0.25, label="Engine"), color="#69b3a2") +
# Bottom
geom_density( aes(x = var2, y = -..density..), fill= "#404080") +
geom_label( aes(x=4.5, y=-0.25, label="Transmission"),
color="#404080") +
theme_ipsum() +
xlab("value of x")
p1
## Warning: The dot-dot notation (`..density..`) was deprecated in ggplot2 3.4.0.
## ℹ Please use `after_stat(density)` instead.
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
## Warning in geom_label(aes(x = 4.5, y = 0.25, label = "Engine"), color = "#69b3a2"): All aesth
etics have length 1, but the data has 32 rows.
## ℹ Please consider using `annotate()` or provide this layer with data containing
## a single row.
## Warning in geom_label(aes(x = 4.5, y = -0.25, label = "Transmission"), color = "#404080"): Al
l aesthetics have length 1, but the data has 32 rows.
## ℹ Please consider using `annotate()` or provide this layer with data containing
## a single row.
## Warning in [Link](C_stringMetric, [Link](x$label)): font family
## not found in Windows font database
## Warning in [Link](C_stringMetric, [Link](x$label)): font family
## not found in Windows font database
[Link] R publish/R pubs/[Link] 3/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
# data Transmisi dan Engine
data <- [Link](
var1 = dd$hp,
var2 = dd$cyl)
p2 <- ggplot(data, aes(x=x) ) +
# Top
geom_density( aes(x = var1, y = ..density..), fill="#19b3a2" ) +
geom_label(aes(x=4.5, y=0.25, label="hp"),
color="#19b3a2") +
# Bottom
geom_density( aes(x = var2, y = -..density..), fill= "#440080") +
geom_label( aes(x=4.5, y=-0.25, label="cyl"),
color="#440080") +
theme_ipsum() +
xlab("value of x")
p2
[Link] R publish/R pubs/[Link] 4/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
## Warning in geom_label(aes(x = 4.5, y = 0.25, label = "hp"), color = "#19b3a2"): All aesthetic
s have length 1, but the data has 32 rows.
## ℹ Please consider using `annotate()` or provide this layer with data containing
## a single row.
## Warning in geom_label(aes(x = 4.5, y = -0.25, label = "cyl"), color = "#440080"): All aesthet
ics have length 1, but the data has 32 rows.
## ℹ Please consider using `annotate()` or provide this layer with data containing
## a single row.
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
2. Violin & Boxplot Chart
[Link] R publish/R pubs/[Link] 5/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
library(dplyr)
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(viridis)
## Loading required package: viridisLite
library(forcats)
library(ggplot2)
mpg %>%
mutate(class = fct_reorder(class, hwy, .fun='length' )) %>%
ggplot( aes(x=class, y=hwy, fill=class)) +
geom_boxplot() +
xlab("class") +
theme([Link]="none") +
xlab("") +
xlab("")
[Link] R publish/R pubs/[Link] 6/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
mpg$class = with(mpg, reorder(class, cty, median))
p2 <- mpg %>%
ggplot( aes(x=class, y=cty, fill=class)) +
geom_violin() +
xlab("class") +
theme([Link]="none") +
xlab("")
p2
[Link] R publish/R pubs/[Link] 7/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
3. Scatter Plot
ggplot(dd, aes(x=hp, y=drat, alpha=am)) +
geom_point(size=6, color="#69b3a2") +
ggtitle("Gross horsepower vs Rear axle ratio") +
theme_ipsum()
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_stringMetric, [Link](x$label)): font family
## not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_stringMetric, [Link](x$label)): font family
## not found in Windows font database
[Link] R publish/R pubs/[Link] 8/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
[Link] R publish/R pubs/[Link] 9/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
ggplot(df, aes(x=hwy, y=cty, shape=drv,
alpha=drv, size=drv, color=drv)) +
ggtitle("City miles/gallon vs Highwaya miles/gallon by Type") +
geom_point() + theme_ipsum()
## Warning: Using alpha for a discrete variable is not advised.
## Warning: Using size for a discrete variable is not advised.
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_textBounds, [Link](x$label), x$x, x$y, : font
## family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
## Warning in [Link](C_text, [Link](x$label), x$x, x$y, :
## font family not found in Windows font database
[Link] R publish/R pubs/[Link] 10/11
8/29/25, 4:31 PM Exploratory Data Aanalysis
[Link] R publish/R pubs/[Link] 11/11