UX Research

Harvard Public Domain Corpus

I led an 8-week research study on the Corpus, delivering insights to drive AI development for researchers across the world

Role

UXR Project Lead

Domain

Service Design

Duration

8 weeks

Team

2 UXR Interns

Harvard Library logo, featuring the Harvard shield crest beside the wordmark.

OVERVIEW

The Corpus is a dataset of 1M books that publicly launched in 2025, but the team didn't have access to data that told them about the tool's impact or performance

Corpus data is prepared as machine-readable text for computational analysis and llm training

Diagram comparing two timelines across the Software Development Lifecycle, showing accessibility currently surfacing during testing versus a future state where it's addressed during design.

My task

Lead research into how users use the Corpus to inform improvement strategies. I owned recruitment, data collection, analysis, and recommendation development

Impact

At the end of my internship, I pitched my recommendations to drive Corpus utilization for current and future users, securing approval from the team and UX & Discovery Director

SERVICE MODEL

The corpus is made available via request and is hosted on an external tool called Globus

Gaining access requires that use be non-profit, educational, or research-related to comply with policy

Diagram comparing two timelines across the Software Development Lifecycle, showing accessibility currently surfacing during testing versus a future state where it's addressed during design.

Objective

Understand how individuals are using corpus data in their work to identify opportunities for improvement

These insights will inform improvement strategies to drive utilization, supporting both new and existing users in their goals of AI research and development

EXISTING DATA

Surprisingly, most users were interested in self-exploration despite the Corpus being designed for llm training

The team suspected users driven by self-exploration may experience high friction because the Corpus was intended for technical use

29%

LLM training

25%

Academic research

45%

Self-exploration & reading

Access requests, n =211

RESEARCH PROCESS

My methodology was centered around 3 core questions

  1. How do users use Public Domain Corpus Data?

  2. What makes it hard or easy to use the corpus data in their work?

  3. What outcomes or outputs has this corpus enabled?

Recruiting for interviews

I sent a survey to all 211 users who requested access to the Corpus both as a way to gather data and recruit respondents for interviews. From 15 respondents, I recruited 2 users and 2 stakeholders for Corpus interviews

Participant tracker
IDRoleUsed Corpus Data?Perspective they giveInterviewer
P1HobbyistNoNon-technical userKelly
P2FacultyYesTechnical userAlly
S1Project LiasonN/ASolution constrainstsAlly
S2Service OwnerN/ARequirementsAlly

Methodology

15 Survey Responses

I sent a survey to all 211 users who accessed the Corpus to identify trends in user engagement and satisfaction

Sample survey questions

4 Interviews

I conducted 2 user interviews to uncover where friction occurs in their journey, and 2 stakeholder interviews to learn about known issues and recommended improvements

Sample interview questions

CONSTRAINTS

Navigating small sample sizes and feasible solutions

  1. Small sample size

From an external pool of potential participants, I received only 2 users who signed up for interviews. I pivoted to conducting stakeholder interviews to move the project along and learn more about known issues and what is feasible for possible solutions

Diagram comparing two timelines across the Software Development Lifecycle, showing accessibility currently surfacing during testing versus a future state where it's addressed during design.
  1. Scope of feasible solutions

Since Globus is an external tool, the feasibility extended to how the data is structured, rather than how the platform is designed

Globus tool

Cloud-based service platform that hosts Corpus, designed to move many large files

Dataset documentation

Additional types of info about the Corpus available to users with access

Library.Harvard homepage

Public-facing webpage where users can learn about the tool and request access

Insight 1

Despite gaining access, most users did not use Corpus data in their work

Diagram comparing two timelines across the Software Development Lifecycle, showing accessibility currently surfacing during testing versus a future state where it's addressed during design.

Those who did use the data are "Expert users," sharing 3 things in common

Expert users

Technical expertise

Storage

Software tools

Non-expert users

Technical expertise

Storage

Software tools

Stakeholders assumed expert users would be the ones who ended up using Corpus data because they had these factors. But these factors were not addressed as requirements or considerations to users, creating friction for non-expert users

Recommendation

Bridge technical knowledge gaps among non-expert users for easier use with links to resources, FAQs, and a persistent contact point

Insight 2

Most participants encountered friction upon initial exploration of the Corpus and quit using it

66%

of survey respondents stopped engaging after briefly exploring the Corpus

of survey respondents stopped engaging after briefly exploring the Corpus

33%

of survey respondents stopped engaging before exploring the data

of survey respondents stopped engaging before exploring the data

Finding

Difficulty understanding dataset quality and lack of time and resources to proceed with their project were among the top cited reasons for abandonment

Recommendation

Set expectations for data quality and technical requirements (e.g. computing capabilities) on the Library.Harvard page

Insight 3

Finding specific files was very difficult even for expert users

"I haven't published anything yet because building the pipeline to extract the metadata to locate books is taking a long time”

Expert user

Finding

Filtering and locating specific books were challenging. Both expert and non-expert users found folder and file names vague

Recommendation

Orient users to the Corpus with documentation on the file structure of the Corpus with with basic details about folders

What this means

Users abandon the Corpus because they don't know what to expect, where to find the data they need, or how to get help

Impact

I pitched my recommendations to drive Corpus utilization for non-expert users, securing approval from the team and UX & Discovery Director

Both expert and non-expert users experience similar friction using the Corpus. Addressing the needs of non-expert users will streamline the workflow of expert users, driving utilization.

Set expectations

surface data quality and technical requirements for transparency

Orient users

document the Corpus's file structure so users can find what they're looking for

Make help easy to find

FAQs, video tutorials, or a persistent contact point for support

" Ally, something I've noticed over the course of the internship is that you are someone who often catch something, the rest of the team might not have considered and the projects comes out better because your questions. Your persistence to learn about the LLM project despite all it's challenges was very impressive. "

Meg McMahon, UX Manager

Takeaways

  1. If I were to continue this project, I would keep recruiting with targeted recruitment with Harvard affiliates until I've reached enough participants (n=15)

  2. Leverage AI to explain technical concepts that appear in interviews and surveys to gain subject matter understanding

  3. Design a survey to still get insights in the case that we do not hear back from many participants

  4. Bring in stakeholders early and often throughout the research process to validate direction