Exploratory Data Analysis on Large Data Sets: The Example of Salary Variation in Spanish Social Security Data

Cookie settings

Necessary

These necessary cookies are required to enable the core functionality of the website. Opting out of these cookies is not possible.

cb-enable

This cookie stores the user's cookie consent status for the current domain. Expiry: 1 year.

laravel_session

Stores the session ID to recognize the user when the page reloads and to restore their login session. Expiry: 2 hours.

XSRF-TOKEN

Provides CSRF protection for forms. Expiry: 2 hours.

Startseite
Publikationen
IZA Discussion Papers
Exploratory Data Analysis on Large Data Sets: The Example of Salary Variation in...

IZA Discussion Paper No. 13459

July 2020

Exploratory Data Analysis on Large Data Sets: The Example of Salary Variation in Spanish Social Security Data

Catia Nicodemo, Albert Satorra

published in: BRQ Business Research Quarterly, 2022, 25 (3), 283–294

New challenges arise in data visualization when a sizable database is used in the analysis. With many data points, classical scatterplots are non-informative due to the cluttering of points. On the contrary, simple plots such as the boxplot that are of limited use in small samples, offer great potential to facilitate group comparison in the case of an extensive sample. This paper presents Exploratory Data Analysis (EDA) methods that are useful when a large dataset is involved. The EDA methods, (introduced by Tukey in his seminal book of 1977) encompass a set of statistical tools aimed to extract information from data using simple graphical tools. In this paper, some of the EDA methods like the Boxplot and Scatterplot are revisited and enhanced using modern graphical computational devices (as, e.g., the heat-map) and their use illustrated with Spanish Social Security data. We explore how earnings vary across several factors like age, gender, type of occupation and contract and in particular, the gender gap in salaries is visualized in various dimensions relating to the type of occupation. The EDA methods are also applied to assessing competing regressions with earnings as the dependent variable. The methods discussed should be useful to researchers to assess heterogeneity in data, across group-variation, and classical diagnostic plots of residuals from alternative models fits.

Download

Keywords

heat-maps EDA Analysis large dataset ggplot R

JEL Codes

C55 J01 J08 Y10 C80