# A tibble: 2,486 × 20
...1 state_territory governor party first_year years_in_office school
<dbl> <chr> <chr> <chr> <dbl> <chr> <chr>
1 0 Alabama Kay Ivey Repu… 2017 2017 - 2019; 2… Unive…
2 1 Alaska Mike Dunleavy Repu… 2018 2018 - 2022; 2… Miser…
3 2 American Samoa Pula’ali’i Nik… Repu… 2025 2025 - 2025 Menlo…
4 3 Arizona Katie Hobbs Demo… 2023 2023 - 2025 North…
5 4 Arkansas Sarah Huckabee… Repu… 2023 2023 - 2025 Ouach…
6 5 California Gavin Newsom Demo… 2019 2019 - 2023; 2… Santa…
7 6 Colorado Jared Polis Demo… 2019 2019 - 2023; 2… Princ…
8 7 Connecticut Ned Lamont Demo… 2019 2019 - 2023; 2… Harva…
9 8 Delaware Matt Meyer Demo… 2025 2025 - 2025 Brown…
10 9 Florida Ron DeSantis Repu… 2019 2019 - 2023; 2… Yale …
# ℹ 2,476 more rows
# ℹ 13 more variables: birth_state_territory <chr>, birth_date <date>,
# bio_text <chr>, college_attendance <dbl>, ivy_attendance <dbl>,
# lawyer <dbl>, military_service <dbl>, age_at_start <dbl>, gender <chr>,
# born_in_state_territory <dbl>, intl_born <dbl>, intl_born_details <chr>,
# `race/ethnicity` <chr>
Project 01
Important dates
- Report: due October 7 at 11:59pm
- Mock interview: October 14-21
Learning objectives
By the end of this project, you will:
- Design a visualization to communicate insights from a dataset
- Create and refine your visualization
- Communicate your design process
- Reflect on your design choices and the effectiveness of your visualization
Introduction
TL;DR: Create a high-quality data visualization and talk about it.
In this project, you will create a data visualization that effectively communicates insights from a dataset provided by the instructor. You will go through the process of exploring the data, sketching and ideating on an initial chart design, creating a rough draft of the chart using R, and refining your visualization to a polished finish. After submitting your visualization, you will participate in a mock interview to discuss your design choices and reflect on the effectiveness of your chart.
Deliverables
The primary deliverables for the project are:
- A report documenting your entire design process
- A mock interview with the instructional staff
Dataset
For the project you will analyze the biographical backgrounds of United States governors (of states and territories) from 1775 to the present.
Accessing the dataset
The dataset is available in the data/ directory of this repository as a CSV file named governor-bios.csv.
What’s in the data?
Alongside the names and states of each governor, the dataset includes information on gender, educational background, as well as other characteristics. The dataset is based on the publicly-available biographies of governors from the National Governors Association (NGA). It contains the following variables:
state_territory: U.S. state or territory where the governor served.governor: Full name of the governor, as listed on the website. The listed name is not always the legal name of the person, e.g. John Ellis Bush Jr. is listed as Jeb Bush.party: Political party affiliation under which the governor served.years_in_office: Concatenated string denoting term(s) in office (e.g., “1985–1991; 1995–1999”). Parsed from web biographies.school: Educational institution(s) attended by the governor as listed on the official profile. If multiple schools are listed, usually separated by semicolons (e.g., “Brigham Young University; Harvard University”). If NA, either no school is listed or the governor did not attend college.birth_state_territory: Parsed location (state) of birth based on text extraction.spouse: Name of spouse, if available. If NA, the governor is either not married or spousal information is unknown.birth_date: Date of birth, as listed on website.bio_text: Raw string of the biographical background of the governor. In 14 cases (less than 1 percent of cases), this biographical text is missing.college_attendance: Equals 1 if the governor attended any college or university; 0 otherwise.ivy_attendance: Equals 1 if the governor attended an Ivy League institution; 0 otherwise. (Note that this technically applies for EITHER graduate or undergraduate).lawyer: Equals 1 if the biography indicates legal education or professional law practice; 0 otherwise.military_service: Equals 1 if the biography mentions military service (e.g., Army, Navy, National Guard, etc.); 0 otherwise.age_at_start: Age of the governor at the beginning of their first term, calculated from the reported date of birth and the start date of their first gubernatorial term.gender: We manually fill in gender based on the Wikipedia list of female governors. All other governors are set to male.born_in_state_territory: 1 if the governor was born in the same U.S. state/territory they governed; 0 if born elsewhere (including outside the U.S.).intl_born: Equals 1 if the governor was born outside of the United States (as defined by the contemporary definition of the 50 states and territories); 0 otherwise. To signify internationally-born, if birth_state was labeled as “Other”, we recognize these governors as international.intl_born_details: To get more information on governors born outside the US, we search the bio text of governors with birth_state equivalent to “Other.” We extract their likely birthplace as a string.race/ethnicity: We use Wikipedia to source a list of all minoritized governors for the states. Then we go through our governors from the territories, and manually checked their race/ethnicity from their Bios and Wikipedia page. We broadly group into the following major categories: White, Latino/Latine/Latinx, Asian American and Pacific Islander, African American, Native American.
The dataset was retrieved from Responsible Datasets in Context and originally collected by White et al. (2026). You can find more information about the dataset, including how it was collected and its limitations, in the original post on the Responsible Datasets in Context website.
Note there are some minimal visualizations included on the website. You are welcome to read the post and these visualizations, but the visualization for your project must be your own original work.
Project workflow
Report
You will create a single polished, high-quality visualization using R and the assigned dataset. The report will be generated using Quarto and rendered as a PDF document to submit via Gradescope.
The report documents your entire design process for the data visualization. It is modeled on Nicola Rennie’s The Art of Data Visualization with ggplot2: The TidyTuesday Cookbook, and should be structured with the following sections and content.
The default Quarto settings for rendering plots in Typst documents are not always ideal for high-detail visualizations. You may want to adjust the default plot dimensions for this project to ensure your visualizations are clear and legible. This tutorial explains in detail how to adjust plot sizing in Quarto documents. In particular, you may want to set the fig-width and fig-height chunk options for your plots to ensure they are large enough to be legible when rendered in the report.
You are responsible for ensuring your plot is clearly legible in the final PDF report.
In the report, make sure to print your code so we can see it in the PDF document. This will facilitate evaluation of your written report and the mock interview. The provided template Quarto document is already configured in the YAML header to print your code:
execute:
echo: trueDataset
- Load the required packages
- Import the dataset
- Briefly describe the dataset, its structure, and its variables
Exploratory work
Learn more about the data and begin to formulate ideas for your visualization.
Data exploration
- Summarize and visualize key aspects of the dataset
- Identify interesting patterns, trends, or relationships in the data
- Document your findings - what are you learning through this exploration? How is this informing your visualization ideas?
Exploratory sketches
- Sketch out at least two distinct visualization ideas on paper or using a digital tool1
- For each sketch, describe:
- The grammar of graphics for the chart (e.g., layers, mappings, scales). It need not be completely worked out yet – for example, you may not have decided on the exact color palette, but you should have a clear idea of the basic structure of the chart and how the data will be mapped to visual channels.
- The rationale behind your design choices (e.g., chart type, color scheme, layout)
- The intended message or insight to be communicated
- Reflect on the strengths and weaknesses of each sketch and explain why you ultimately chose one to pursue
There are lots of guides and tools online about how to select an appropriate chart type based on the data you have and the message you want to communicate. I personally like From Data to Viz, but there are many others out there.
Preparing a plot
Begin creating your visualization in R based on your chosen sketch.
Data wrangling
- Perform any required data wrangling to prepare the dataset for visualization
- Document the steps taken and explain why they were necessary for your visualization
The first plot
- Create a functional first draft of your chosen visualization
- It need not be polished or final, but it should convey the basic structure and message of your intended chart
- Essentially it should have all the grammatical components of the chart (e.g. layers, mappings, scales, etc.) but you do not need to have any of the styling or theming worked out yet
Advanced styling
Make it shine! This is where you take the basic plot you created in the previous section and refine it to a polished, high-quality visualization. Adjustments you will likely make include:
- Fine-tuning colors, fonts, and other stylistic elements
- Adding titles, labels, and annotations to enhance clarity
- Implementing custom themes or styles to align with your design vision
- Improving layout, spacing, and aspect ratio for better readability
Reflection
Given the time constraints, it’s unlikely that your chart will be perfect. However, it’s important to reflect on your design choices and the effectiveness of your visualization. Address the following questions in your reflection:
- How well does your final visualization communicate the intended message or insight?
- What design choices did you make to enhance clarity and engagement?
- What challenges did you encounter during the design process, and how did you address them?
- If you had more time, what additional improvements or refinements would you make to your visualization?
Generative AI (GAI) self-reflection
As stated in the syllabus, include a written reflection for this assignment of how you used GAI tools (e.g. what tools you used, how you used them to assist you with writing code), what skills you believe you acquired, and how you believe you demonstrated mastery of the learning objectives.
Mock interview
Each student will participate in a 15 minute mock interview. During the mock interview, you will discuss your design process and answer questions about your visualization. The interview will cover topics such as:
- Your data exploration and insights
- The rationale behind your chosen visualization design
- Specific design choices and their intended effects
- Reflections on the effectiveness of your visualization
- Implementation details and your code
Wrap up
Submission
Report
- Go to http://www.gradescope.com and click Log in in the top right corner.
- Click School Credentials \(\rightarrow\) Cornell University NetID and log in using your NetID credentials.
- Click on your INFO 3312 course.
- Click on the assignment, and you’ll be prompted to submit it.
- Mark all the pages for the assignment as Report and click Submit.
Mock interview
Sign up for an interview time slot to be held between October 14-21.
Grading and evaluation criteria
| Total | 100 pts |
|---|---|
| Report | 50 pts |
| Mock interview | 50 pts |
Report
| Category | Less developed projects | Typical projects | More developed projects |
|---|---|---|---|
| Data exploration + insight | Exploration is superficial or disconnected from visualization choices. Findings are not clearly articulated or don’t motivate the final design. | Thorough exploration reveals key patterns and relationships. Clear documentation of findings that directly inform visualization choices. Shows genuine discovery process. | All expectations of typical projects + exploration uncovers non-obvious or compelling insights. Articulates a clear, focused narrative that the visualization will communicate. |
| Design thinking + justification | Sketches are missing or lack description. Rationale for design choices is absent or poorly explained. Little evidence of deliberate decision-making about chart type, variables, or visual encodings. | Presents multiple sketch ideas with clear descriptions of chart type, variables, and design rationale. Explains why the chosen design is appropriate for the data and message. Reflects on trade-offs between options. | All expectations of typical projects + sketches demonstrate sophisticated thinking about design alternatives and how different designs communicate different messages. Shows deep consideration of visual hierarchy, data-ink ratio, and other design principles. |
| Chart type + grammar | Chart type is inappropriate for the data or message. Missing or incorrect mappings of variables to visual channels. Basic grammatical components (layers, scales) are incomplete or incorrect. | Chart type is appropriate and well-justified. Variables are effectively mapped to visual channels. All grammatical components (layers, mappings, scales) are present and correct. Code is functional and readable. | All expectations of typical projects + uses sophisticated or layered designs (e.g., faceting, multiple geoms) where appropriate. Demonstrates mastery of the grammar of graphics and intentional use of visual encoding to strengthen the message. |
| Visual design | Visualization is difficult to read. Poor color choices, illegible fonts, or confusing layout. Lacks labels or annotations. Does not follow best practices. | Visualization is polished and easy to read. Appropriate color palette with sufficient contrast. Clean typography and well-organized layout. Clear titles, axis labels, and legends. Follows visualization best practices taught in class. | All expectations of typical projects + employs sophisticated visual design with custom themes, distinctive color schemes, or effective use of whitespace and visual hierarchy. Color choices show understanding of colorblindness accessibility. Typography and styling enhance clarity and engagement. |
| Reflection | Reflection is missing or superficial. Does not meaningfully address design choices or effectiveness. | Reflection addresses design choices and their effects. Shows honest assessment of what worked and what didn’t. Demonstrates understanding of how visualization choices communicate (or fail to communicate) the intended message. | All expectations of typical projects + reflection demonstrates sophisticated understanding of why certain designs are more effective. Identifies specific, actionable improvements. Shows evidence of iterative thinking and learning throughout the design process. |
Mock interview
All students will be evaluated on a standardized rubric which will be provided to students after the exam. You should be prepared to discuss any of the design and/or implementation choices in your project.
Late work policy
There is no late work accepted on this project. Your report must be submitted by the deadline.
Students are expected to attend their scheduled interview time slot. If you miss your interview without advance notice or sufficient rationale, you will earn a 20% deduction on your mock interview score. If you know in advance that you will miss your scheduled interview, please contact the instructor as soon as possible to discuss accommodations.
References
Footnotes
If sketched on paper, take a legible photo and include it in your report. Do not submit Canva/Figma wireframe mockups. It must be drawn free form.↩︎