0% found this document useful (0 votes)

61 views19 pages

Data - Preprocessing 1 19

Chapter 3 discusses data preprocessing, emphasizing the importance of data quality and the major tasks involved, including data cleaning, integration, reduction, and transformation. It outlines the challenges of handling missing, noisy, and inconsistent data, as well as techniques for data cleaning and integration. The chapter also covers methods for evaluating data quality and resolving conflicts during data integration.

Uploaded by

Mahim Jain Anwa

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

61 views19 pages

Data - Preprocessing 1 19

Uploaded by

Mahim Jain Anwa

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 19

Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
1
Data Quality: Why Preprocess the Data?

◼ Measures for data quality: A multidimensional view

◼ Accuracy: correct or wrong, accurate or not
◼ Completeness: not recorded, unavailable, …
◼ Consistency: some modified but some not, dangling, …
◼ Timeliness: timely update?
◼ Believability: how trustable the data are correct?
◼ Interpretability: how easily the data can be
understood?

2
Major Tasks in Data Preprocessing
◼ Data cleaning
◼ Fill in missing values, smooth noisy data, identify or remove
outliers, and resolve inconsistencies
◼ Data integration
◼ Integration of multiple databases, data cubes, or files
◼ Data reduction
◼ Dimensionality reduction
◼ Numerosity reduction
◼ Data compression
◼ Data transformation and data discretization
◼ Normalization
◼ Concept hierarchy generation

3
Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
4
Data Cleaning
◼ Data in the Real World Is Dirty: Lots of potentially incorrect data,
e.g., instrument faulty, human or computer error, transmission error
◼ incomplete: lacking attribute values, lacking certain attributes of
interest, or containing only aggregate data
◼ e.g., Occupation=“ ” (missing data)
◼ noisy: containing noise, errors, or outliers
◼ e.g., Salary=“−10” (an error)
◼ inconsistent: containing discrepancies in codes or names, e.g.,
◼ Age=“42”, Birthday=“03/07/2010”
◼ Was rating “1, 2, 3”, now rating “A, B, C”
◼ discrepancy between duplicate records
◼ Intentional (e.g., disguised missing data)
◼ Jan. 1 as everyone’s birthday?
5
Incomplete (Missing) Data

◼ Data is not always available

◼ E.g., many tuples have no recorded value for several
attributes, such as customer income in sales data
◼ Missing data may be due to
◼ equipment malfunction
◼ inconsistent with other recorded data and thus deleted
◼ data not entered due to misunderstanding
◼ certain data may not be considered important at the
time of entry
◼ not register history or changes of the data
◼ Missing data may need to be inferred
6
How to Handle Missing Data?
◼ Ignore the tuple: usually done when class label is missing
(when doing classification)—not effective when the % of
missing values per attribute varies considerably
◼ Fill in the missing value manually: tedious + infeasible?
◼ Fill in it automatically with
◼ a global constant : e.g., “unknown”, a new class?!
◼ the attribute mean
◼ the attribute mean for all samples belonging to the
same class: smarter
◼ the most probable value: inference-based such as
Bayesian formula or decision tree
7
Noisy Data
◼ Noise: random error or variance in a measured variable
◼ Incorrect attribute values may be due to
◼ faulty data collection instruments

◼ data entry problems

◼ data transmission problems

◼ technology limitation

◼ inconsistency in naming convention

◼ Other data problems which require data cleaning

◼ duplicate records

◼ incomplete data

◼ inconsistent data

8
How to Handle Noisy Data?

◼ Binning
◼ first sort data and partition into (equal-frequency) bins

◼ then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

◼ Regression
◼ smooth by fitting the data into regression functions

◼ Clustering
◼ detect and remove outliers

◼ Combined computer and human inspection

◼ detect suspicious values and check by human (e.g.,

deal with possible outliers)

9
Data Cleaning as a Process
◼ Data discrepancy detection
◼ Use metadata (e.g., domain, range, dependency, distribution)

◼ Check field overloading

◼ Check uniqueness rule, consecutive rule and null rule

◼ Use commercial tools

◼ Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

◼ Data auditing: by analyzing data to discover rules and

relationship to detect violators (e.g., correlation and clustering

to find outliers)
◼ Data migration and integration
◼ Data migration tools: allow transformations to be specified

◼ ETL (Extraction/Transformation/Loading) tools: allow users to

specify transformations through a graphical user interface
◼ Integration of the two processes
◼ Iterative and interactive (e.g., Potter’s Wheels)

10
Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
11
Data Integration
◼ Data integration:
◼ Combines data from multiple sources into a coherent store
◼ Schema integration: e.g., A.cust-id  B.cust-#
◼ Integrate metadata from different sources
◼ Entity identification problem:
◼ Identify real world entities from multiple data sources, e.g., Bill
Clinton = William Clinton
◼ Detecting and resolving data value conflicts
◼ For the same real world entity, attribute values from different
sources are different
◼ Possible reasons: different representations, different scales, e.g.,
metric vs. British units
12
Handling Redundancy in Data Integration

◼ Redundant data occur often when integration of multiple

databases
◼ Object identification: The same attribute or object
may have different names in different databases
◼ Derivable data: One attribute may be a “derived”
attribute in another table, e.g., annual revenue
◼ Redundant attributes may be able to be detected by
correlation analysis and covariance analysis
◼ Careful integration of the data from multiple sources may
help reduce/avoid redundancies and inconsistencies and
improve mining speed and quality
13
Correlation Analysis (Numeric Data)

◼ Correlation coefficient (also called Pearson’s product

moment coefficient)

i=1 (ai − A)(bi − B) 

n n
(ai bi ) − n AB
rA, B = = i =1
(n − 1) A B (n − 1) A B

where n is the number of tuples, A and B are the respective

means of A and B, σA and σB are the respective standard deviation
of A and B, and Σ(aibi) is the sum of the AB cross-product.
◼ If rA,B > 0, A and B are positively correlated (A’s values
increase as B’s).
◼ rA,B = 0: independent; rAB < 0: negatively correlated

14
2/12/2025 Data Mining: Concepts and Techniques 15
Visually Evaluating Correlation

Scatter plots
showing the
similarity from
–1 to 1.

16
Correlation (viewed as linear relationship)
◼ Correlation measures the linear relationship
between objects
◼ To compute correlation, we standardize data
objects, A and B, and then take their dot product

a 'k = (ak − mean( A)) / std ( A)

b'k = (bk − mean( B )) / std ( B)

correlation( A, B) = A'• B '

17
Covariance (Numeric Data)
◼ Covariance is similar to correlation

Correlation coefficient:

where n is the number of tuples, A and B are the respective mean or

expected values of A and B, σA and σB are the respective standard
deviation of A and B.
◼ Positive covariance: If CovA,B > 0, then A and B both tend to be larger
than their expected values.
◼ Negative covariance: If CovA,B < 0 then if A is larger than its expected
value, B is likely to be smaller than its expected value.
◼ Independence: CovA,B = 0 but the converse is not true:
◼ Some pairs of random variables may have a covariance of 0 but are not
independent. Only under some additional assumptions (e.g., the data follow
multivariate normal distributions) does a covariance of 0 imply independence18
Co-Variance: An Example

◼ It can be simplified in computation as

◼ Suppose two stocks A and B have the following values in one week:
(2, 5), (3, 8), (5, 10), (4, 11), (6, 14).

◼ Question: If the stocks are affected by the same industry trends, will
their prices rise or fall together?

◼ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4

◼ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6

◼ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4

◼ Thus, A and B rise together since Cov(A, B) > 0.

Chapter 3 - Tagged
No ratings yet
Chapter 3 - Tagged
63 pages
Wk6 Preprocessing
No ratings yet
Wk6 Preprocessing
64 pages
DM Merged
No ratings yet
DM Merged
169 pages
Lec 7
No ratings yet
Lec 7
45 pages
Unit 1 C
No ratings yet
Unit 1 C
63 pages
Data Quality and Preprocessing Techniques
No ratings yet
Data Quality and Preprocessing Techniques
63 pages
03preprocessing 20160222
No ratings yet
03preprocessing 20160222
65 pages
Data Preprocessing for Regression Analysis
No ratings yet
Data Preprocessing for Regression Analysis
56 pages
Chapter 3
No ratings yet
Chapter 3
63 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
63 pages
03 Preprocessing
No ratings yet
03 Preprocessing
63 pages
Data Preprocessing
No ratings yet
Data Preprocessing
63 pages
IT446 Wk03.2 HanKamberPei 03preprocessing PDF
No ratings yet
IT446 Wk03.2 HanKamberPei 03preprocessing PDF
64 pages
Data Preprocessing: Discretization Techniques
No ratings yet
Data Preprocessing: Discretization Techniques
63 pages
Mining
No ratings yet
Mining
63 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
64 pages
03 Preprocessing
No ratings yet
03 Preprocessing
54 pages
Data Preprocessing
No ratings yet
Data Preprocessing
77 pages
03 Preprocessing
No ratings yet
03 Preprocessing
65 pages
Module 2
No ratings yet
Module 2
62 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
66 pages
Data Mining: Dosen: Dr. Vitri Tundjungsari
No ratings yet
Data Mining: Dosen: Dr. Vitri Tundjungsari
64 pages
Data Pre Processing
No ratings yet
Data Pre Processing
63 pages
Module 5 03preprocessing
No ratings yet
Module 5 03preprocessing
63 pages
03 Preprocessing
No ratings yet
03 Preprocessing
64 pages
Lecture 2.3.1-2.3.3
No ratings yet
Lecture 2.3.1-2.3.3
67 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
61 pages
03 Preprocessing
No ratings yet
03 Preprocessing
60 pages
Data Preprocessing Essentials
No ratings yet
Data Preprocessing Essentials
56 pages
03 Pre Processing
No ratings yet
03 Pre Processing
63 pages
Data Preprocessing Techniques Overview
No ratings yet
Data Preprocessing Techniques Overview
65 pages
Preprocessing Techniques
No ratings yet
Preprocessing Techniques
63 pages
03 Pre Processing
No ratings yet
03 Pre Processing
89 pages
Lecture 3
No ratings yet
Lecture 3
47 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
62 pages
Data Preprocessing Techniques
No ratings yet
Data Preprocessing Techniques
52 pages
Unit 2 Data Preprocessing
No ratings yet
Unit 2 Data Preprocessing
40 pages
Lecture#2 Data Mining MS (DEIM) Spring 2025
No ratings yet
Lecture#2 Data Mining MS (DEIM) Spring 2025
61 pages
Data Mining 3
No ratings yet
Data Mining 3
57 pages
03 Preprocessing
No ratings yet
03 Preprocessing
38 pages
Slide 05 Chapter3 Data Preprocessing
No ratings yet
Slide 05 Chapter3 Data Preprocessing
58 pages
Data Preprocessing Overview and Techniques
100% (1)
Data Preprocessing Overview and Techniques
41 pages
Lec 3
No ratings yet
Lec 3
31 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
54 pages
Data Preprocessing (Sagar)
No ratings yet
Data Preprocessing (Sagar)
31 pages
Unit2 Part2
No ratings yet
Unit2 Part2
67 pages
Data Pre Processing
No ratings yet
Data Pre Processing
62 pages
Module 2 (C) - Data Preprocessing
No ratings yet
Module 2 (C) - Data Preprocessing
50 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
30 pages
PPT1
No ratings yet
PPT1
93 pages
2020 Preprocessing
No ratings yet
2020 Preprocessing
63 pages
Major Tasks in Data Preprocessing
No ratings yet
Major Tasks in Data Preprocessing
62 pages
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
No ratings yet
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
37 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
50 pages
Unit 3
No ratings yet
Unit 3
164 pages
3 Processing
No ratings yet
3 Processing
79 pages
DataScience 1
No ratings yet
DataScience 1
22 pages
Oec Project Final
No ratings yet
Oec Project Final
9 pages
Titanic Survival Analysis by Class & Gender
No ratings yet
Titanic Survival Analysis by Class & Gender
2 pages
Assignment
No ratings yet
Assignment
3 pages
Adverse Reactions - CSV
No ratings yet
Adverse Reactions - CSV
1 page
Assignment 8 Report
No ratings yet
Assignment 8 Report
13 pages
Quantum Theory
No ratings yet
Quantum Theory
65 pages
4 EL EMWaves Vaccum BC
No ratings yet
4 EL EMWaves Vaccum BC
46 pages
2 EL Div Curl EB
No ratings yet
2 EL Div Curl EB
50 pages
1 EL Vectors
No ratings yet
1 EL Vectors
59 pages
3 EL Maxwells Eqns
No ratings yet
3 EL Maxwells Eqns
23 pages
Oracle
No ratings yet
Oracle
8 pages
OracleApps88 - Oracle Alerts PDF
No ratings yet
OracleApps88 - Oracle Alerts PDF
13 pages
MVA Implementing A Data Warehouse With SQL Jump Start Mod 1 Final
No ratings yet
MVA Implementing A Data Warehouse With SQL Jump Start Mod 1 Final
37 pages
Neo4j: What's A Graph Database?
No ratings yet
Neo4j: What's A Graph Database?
2 pages
SPPU SYBSc CS Sem3 Theory Papers 2025
No ratings yet
SPPU SYBSc CS Sem3 Theory Papers 2025
5 pages
Untitled
No ratings yet
Untitled
1,984 pages
Changing An Idoc's Status With An Excel Upload: Main Program
0% (1)
Changing An Idoc's Status With An Excel Upload: Main Program
5 pages
The Future of MySQL (The Project)
100% (10)
The Future of MySQL (The Project)
20 pages
05-BI Framework and Components
No ratings yet
05-BI Framework and Components
22 pages
ER Model and Database Design
No ratings yet
ER Model and Database Design
40 pages
1) Oracle Administrator Question & Answers
No ratings yet
1) Oracle Administrator Question & Answers
8 pages
Big Data Tools and Applications Overview
No ratings yet
Big Data Tools and Applications Overview
10 pages
Ruby Calculator User Guide
No ratings yet
Ruby Calculator User Guide
17 pages
CST121 Database Management With Access-Week4, 5 6
No ratings yet
CST121 Database Management With Access-Week4, 5 6
61 pages
DATASET Configuration
No ratings yet
DATASET Configuration
27 pages
Access Doc 1 Notes
100% (1)
Access Doc 1 Notes
50 pages
Big Data Analytics - Notes
No ratings yet
Big Data Analytics - Notes
13 pages
SQL Server Database Development Best Practices: Grant Fritchey, Red Gate Software
No ratings yet
SQL Server Database Development Best Practices: Grant Fritchey, Red Gate Software
21 pages
Assignment Assistance - DBI202 - thuyenPTL - Spring24
No ratings yet
Assignment Assistance - DBI202 - thuyenPTL - Spring24
4 pages
2018 Book NetworkDataAnalytics PDF
100% (1)
2018 Book NetworkDataAnalytics PDF
406 pages
Unit 5
No ratings yet
Unit 5
3 pages
Unit 4
No ratings yet
Unit 4
29 pages
Projek F1082 F1070
No ratings yet
Projek F1082 F1070
22 pages
Project On "Fee Management" By: Sanjeev Bhadauria (PGT CS) KV BARABANKI (Lucknow Region)
No ratings yet
Project On "Fee Management" By: Sanjeev Bhadauria (PGT CS) KV BARABANKI (Lucknow Region)
6 pages
INE Advanced Injection Attacks Course File
No ratings yet
INE Advanced Injection Attacks Course File
258 pages
Top AWS Tools for Machine Learning
No ratings yet
Top AWS Tools for Machine Learning
2 pages
Dbms Notes
No ratings yet
Dbms Notes
46 pages
C# Crystal Reports with SQL Queries
No ratings yet
C# Crystal Reports with SQL Queries
21 pages
Data Engineering Brochure
No ratings yet
Data Engineering Brochure
24 pages
Search and Meta Search Engines
No ratings yet
Search and Meta Search Engines
9 pages

Data - Preprocessing 1 19

Uploaded by

Data - Preprocessing 1 19

Uploaded by

Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Measures for data quality: A multidimensional view

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Data is not always available

◼ data entry problems

◼ data transmission problems

◼ inconsistency in naming convention

◼ Other data problems which require data cleaning

◼ then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

◼ Combined computer and human inspection

deal with possible outliers)

◼ Check field overloading

◼ Check uniqueness rule, consecutive rule and null rule

◼ Use commercial tools

◼ Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

relationship to detect violators (e.g., correlation and clustering

◼ ETL (Extraction/Transformation/Loading) tools: allow users to

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Redundant data occur often when integration of multiple

◼ Correlation coefficient (also called Pearson’s product

i=1 (ai − A)(bi − B) 

where n is the number of tuples, A and B are the respective

a 'k = (ak − mean( A)) / std ( A)

b'k = (bk − mean( B )) / std ( B)

correlation( A, B) = A'• B '

where n is the number of tuples, A and B are the respective mean or

◼ It can be simplified in computation as

◼ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4

◼ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6

◼ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4

◼ Thus, A and B rise together since Cov(A, B) > 0.

You might also like