Introduction to OpenRefine

Session Schedule:

Timed entries are shown in your local timezone.

Description

Researchers and librarians often deal with “messy” data. This could include inconsistent interview transcripts, poorly formatted longitudinal datasets, or standardizing names for an authority file. Preparing this data for analysis is a critical step that requires the same level of rigor and reproducibility as the analysis itself.

This workshop introduces OpenRefine, a powerful, free, and open-source tool for cleaning, normalizing, and transforming research data. participants work with a sample dataset to learn how to efficiently handle common data issues, such as automating data standardization, identifying clusters of similar entries, and transforming complex strings into structured information without the need for advanced programming. Unlike manual editing in Excel, OpenRefine records every step of your process, ensuring your data cleaning is transparent and fully reproducible.

Prerequisites

Participants should be comfortable using a computer and have a general understanding of spreadsheets. No prior programming experience is necessary.

Details
Format
Online
Location
Online
Level
Beginner
Duration
3 hours
Credential
None
Cost
FREE
Learning Objectives

- Differentiate how OpenRefine differs from spreadsheet software like Excel and where it fits in the data lifecycle.
- Configure and filter data by using facets (text, numeric, and timeline) to identify patterns, inconsistencies, and errors within a dataset.
- Deploy OpenRefine’s built-in clustering algorithms to automatically group and resolve spelling variations or typos (e.g., merging "New York," "new york," and "N.Y.").
- Perform basic data transformations to split, merge, and reformat columns to meet metadata standards.
- Optimize reproducibility by using the "Undo/Redo" history to track every change made to a dataset and export those steps to apply to future projects.

Who Should Enrol

Researchers or librarians who work with “messy” data: large volumes of bibliographic or metadata records, inconsistent naming conventions, or the need to standardize data.