Adjusting scales for World Bank indicators
Suggested answers
Data: World economic measures
The World Bank publishes a rich and detailed set of socioeconomic indicators spanning several decades and dozens of topics. Here we focus on a few key indicators for the year 2023.
-
gdp_per_cap- GDP per capita (current USD) -
pop- Total population -
life_exp- Life expectancy at birth, total (years) -
female_labor_pct- Labor force, female (% of total labor force) -
income_level- Classification of economies based on national income levels
The data is stored in wb-indicators.rds. To import the data, use the read_rds() function.
world_bank <- read_rds("data/wb-indicators.rds")Customize scales
Let’s consider the relationship between female labor participation and per capita GDP. We’ll use the income_level variable to color the points and provide context on the overall wealth of the countries.1
Step 1: Base plot
First, let’s generate a color-coded scatterplot with a single smoothing line.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level)) +
geom_smooth(se = FALSE)Step 2: Clear up the axes
Now, let’s modify the scales to make the chart more readable. Log-transform the \(y\)-axis and format the labels so they are explicitly identified as percentages and currency.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level)) +
geom_smooth(se = FALSE) +
scale_x_continuous(labels = label_percent(scale = 1)) +
scale_y_log10(labels = label_currency(scale_cut = cut_short_scale()))Step 3: Label the chart
Add human-readable labels for the title, axes, and legend.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level)) +
geom_smooth(se = FALSE) +
scale_x_continuous(labels = label_percent(scale = 1)) +
scale_y_log10(labels = label_currency(scale_cut = cut_short_scale())) +
labs(
title = "Female labor participation is weakly correlated with per capita GDP",
x = "Female labor (percentage of total workforce)",
y = "GDP per capita (current USD)",
color = "Level of income"
)Step 4: Use an optimal color palette
Use the {viridis} color palette for income_level.
The bright yellow at the end of the palette is hard on the eyes. You can condense the hue at which the color map ends using the end argument to the appropriate scale_color_*() function.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level)) +
geom_smooth(se = FALSE) +
scale_x_continuous(labels = label_percent(scale = 1)) +
scale_y_log10(labels = label_currency(scale_cut = cut_short_scale())) +
scale_color_viridis_d(end = 0.8) +
labs(
title = "Female labor participation is weakly correlated with per capita GDP",
x = "Female labor (percentage of total workforce)",
y = "GDP per capita (current USD)",
color = "Level of income"
)Step 5: Double-encode the income_level variable
Double-encode the income_level variable by using both color and shape to represent the same variable. Condense the guides so you use a single legend.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level, shape = income_level)) +
geom_smooth(se = FALSE) +
scale_x_continuous(labels = label_percent(scale = 1)) +
scale_y_log10(labels = label_currency(scale_cut = cut_short_scale())) +
scale_color_viridis_d(end = 0.8) +
labs(
title = "Female labor participation is weakly correlated with per capita GDP",
x = "Female labor (percentage of total workforce)",
y = "GDP per capita (current USD)",
color = "Level of income",
shape = "Level of income"
)Step 6: Reverse the legend entries
It’s annoying that the order of the values in the legend are opposite from how the income levels are ordered in the chart. Reverse the order of the values in the legend so they correspond to the ordering on the \(y\)-axis.
ggplot(
data = world_bank,
mapping = aes(x = female_labor_pct, y = gdp_per_cap)
) +
geom_point(mapping = aes(color = income_level, shape = income_level)) +
geom_smooth(se = FALSE) +
scale_x_continuous(labels = label_percent(scale = 1)) +
scale_y_log10(labels = label_currency(scale_cut = cut_short_scale())) +
scale_color_viridis_d(end = 0.8, guide = guide_legend(reverse = TRUE)) +
scale_shape_discrete(guide = guide_legend(reverse = TRUE)) +
labs(
title = "Female labor participation is weakly correlated with per capita GDP",
x = "Female labor (percentage of total workforce)",
y = "GDP per capita (current USD)",
color = "Level of income",
shape = "Level of income"
)sessioninfo::session_info()─ Session info ───────────────────────────────────────────────────────────────
setting value
version R version 4.6.1 (2026-06-24)
os macOS Tahoe 26.6.2
system aarch64, darwin23
ui X11
language (EN)
collate en_US.UTF-8
ctype en_US.UTF-8
tz America/New_York
date 2026-09-10
pandoc 3.10 @ /Applications/Positron.app/Contents/Resources/app/quarto/bin/tools/aarch64/ (via rmarkdown)
quarto 1.10.18 @ /Applications/quarto/bin/quarto
─ Packages ───────────────────────────────────────────────────────────────────
! package * version date (UTC) lib source
P cli 3.6.6 2026-04-09 [?] RSPM
P digest 0.6.39 2025-11-19 [?] RSPM
P dplyr * 1.2.1 2026-04-03 [?] RSPM
P evaluate 1.0.5 2025-08-27 [?] RSPM
P farver 2.1.2 2024-05-13 [?] RSPM
P fastmap 1.2.0 2024-05-15 [?] RSPM
P forcats * 1.0.1 2025-09-25 [?] RSPM
P generics 0.1.4 2025-05-09 [?] RSPM
P ggplot2 * 4.0.3 2026-04-22 [?] RSPM
P glue 1.8.1 2026-04-17 [?] RSPM
P gridExtra 2.3.1 2026-06-25 [?] RSPM
P gtable 0.3.6 2024-10-25 [?] RSPM
P here 1.0.2 2025-09-15 [?] RSPM
P hms 1.1.4 2025-10-17 [?] RSPM
P htmltools 0.5.9 2025-12-04 [?] RSPM
P htmlwidgets 1.6.4 2023-12-06 [?] RSPM
P jsonlite 2.0.0 2025-03-27 [?] RSPM
P knitr 1.51 2025-12-20 [?] RSPM
P labeling 0.4.3 2023-08-29 [?] RSPM
P lattice 0.22-9 2026-02-09 [?] CRAN (R 4.6.1)
P lifecycle 1.0.5 2026-01-08 [?] RSPM
P lubridate * 1.9.5 2026-02-04 [?] RSPM
P magrittr 2.0.5 2026-04-04 [?] RSPM
P Matrix 1.7-5 2026-03-21 [?] CRAN (R 4.6.1)
P mgcv 1.9-4 2025-11-07 [?] CRAN (R 4.6.1)
P nlme 3.1-169 2026-03-27 [?] CRAN (R 4.6.1)
P otel 0.2.0 2025-08-29 [?] RSPM
P pillar 1.11.1 2025-09-17 [?] RSPM
P pkgconfig 2.0.3 2019-09-22 [?] RSPM
P purrr * 1.2.2 2026-04-10 [?] RSPM
P R6 2.6.1 2025-02-15 [?] RSPM
P RColorBrewer 1.1-3 2022-04-03 [?] RSPM
P readr * 2.2.0 2026-02-19 [?] RSPM
P renv 1.2.2 2026-04-16 [?] RSPM
P rlang 1.3.0 2026-07-05 [?] RSPM
P rmarkdown 2.31 2026-03-26 [?] RSPM
P rprojroot 2.1.1 2025-08-26 [?] RSPM
P S7 0.2.2 2026-04-22 [?] RSPM
P scales * 1.4.0 2025-04-24 [?] RSPM
P sessioninfo 1.2.4 2026-06-04 [?] RSPM
P stringi 1.8.9 2026-08-04 [?] RSPM
P stringr * 1.6.0 2025-11-04 [?] RSPM
P tibble * 3.3.1 2026-01-11 [?] RSPM
P tidyr * 1.3.2 2025-12-19 [?] RSPM
P tidyselect 1.2.1 2024-03-11 [?] RSPM
P tidyverse * 2.0.0 2023-02-22 [?] RSPM
P timechange 0.4.0 2026-01-29 [?] RSPM
P tzdb 0.5.0 2025-03-15 [?] RSPM
P vctrs 0.7.3 2026-04-11 [?] RSPM
P viridis * 0.6.5 2024-01-29 [?] RSPM
P viridisLite * 0.4.3 2026-02-04 [?] RSPM
P withr 3.0.3 2026-06-19 [?] RSPM
P xfun 0.60 2026-07-09 [?] RSPM
P yaml 2.3.12 2025-12-10 [?] RSPM
[1] /Users/bcs88/Projects/info-3312/course-site/renv/library/macos/R-4.6/aarch64-apple-darwin23
[2] /Users/bcs88/Library/Caches/org.R-project.R/R/renv/sandbox/macos/R-4.6/aarch64-apple-darwin23/46003b10
* ── Packages attached to the search path.
P ── Loaded and on-disk path mismatch.
──────────────────────────────────────────────────────────────────────────────
Footnotes
Note that the income level is based on the GNI per capita, which is strongly correlated with GDP per capita, but not exactly the same.↩︎





