UX Research
Harvard Public Domain Corpus
I led an 8-week research study on the Corpus, delivering insights to drive AI development for researchers across the world
Role
UXR Project Lead
Domain
Service Design
Duration
8 weeks
Team
2 UXR Interns

OVERVIEW
The Corpus is a dataset of 1M books that publicly launched in 2025, but the team didn't have access to data that told them about the tool's impact or performance
Corpus data is prepared as machine-readable text for computational analysis and llm training

My task
Lead research into how users use the Corpus to inform improvement strategies. I owned recruitment, data collection, analysis, and recommendation development
Impact
At the end of my internship, I pitched my recommendations to drive Corpus utilization for current and future users, securing approval from the team and UX & Discovery Director
SERVICE MODEL
The corpus is made available via request and is hosted on an external tool called Globus
Gaining access requires that use be non-profit, educational, or research-related to comply with policy

Objective
Understand how individuals are using corpus data in their work to identify opportunities for improvement
These insights will inform improvement strategies to drive utilization, supporting both new and existing users in their goals of AI research and development
EXISTING DATA
Surprisingly, most users were interested in self-exploration despite the Corpus being designed for llm training
The team suspected users driven by self-exploration may experience high friction because the Corpus was intended for technical use
29%
LLM training
25%
Academic research
45%
Self-exploration & reading
Access requests, n =211
RESEARCH PROCESS
My methodology was centered around 3 core questions
How do users use Public Domain Corpus Data?
What makes it hard or easy to use the corpus data in their work?
What outcomes or outputs has this corpus enabled?
Recruiting for interviews
I sent a survey to all 211 users who requested access to the Corpus both as a way to gather data and recruit respondents for interviews. From 15 respondents, I recruited 2 users and 2 stakeholders for Corpus interviews
| ID | Role | Used Corpus Data? | Perspective they give | Interviewer |
|---|---|---|---|---|
| P1 | Hobbyist | No | Non-technical user | Kelly |
| P2 | Faculty | Yes | Technical user | Ally |
| S1 | Project Liason | N/A | Solution constrainsts | Ally |
| S2 | Service Owner | N/A | Requirements | Ally |
Methodology
15 Survey Responses
I sent a survey to all 211 users who accessed the Corpus to identify trends in user engagement and satisfaction
Sample survey questions
4 Interviews
I conducted 2 user interviews to uncover where friction occurs in their journey, and 2 stakeholder interviews to learn about known issues and recommended improvements
Sample interview questions
CONSTRAINTS
Navigating small sample sizes and feasible solutions
Small sample size
From an external pool of potential participants, I received only 2 users who signed up for interviews. I pivoted to conducting stakeholder interviews to move the project along and learn more about known issues and what is feasible for possible solutions

Scope of feasible solutions
Since Globus is an external tool, the feasibility extended to how the data is structured, rather than how the platform is designed
Globus tool
Cloud-based service platform that hosts Corpus, designed to move many large files
Dataset documentation
Additional types of info about the Corpus available to users with access
Library.Harvard homepage
Public-facing webpage where users can learn about the tool and request access
Insight 1
Despite gaining access, most users did not use Corpus data in their work

Those who did use the data are "Expert users," sharing 3 things in common
Expert users
Technical expertise
Storage
Software tools
Non-expert users
Technical expertise
Storage
Software tools
Stakeholders assumed expert users would be the ones who ended up using Corpus data because they had these factors. But these factors were not addressed as requirements or considerations to users, creating friction for non-expert users
Recommendation
Bridge technical knowledge gaps among non-expert users for easier use with links to resources, FAQs, and a persistent contact point
Insight 2
Most participants encountered friction upon initial exploration of the Corpus and quit using it
66%
33%
Finding
Difficulty understanding dataset quality and lack of time and resources to proceed with their project were among the top cited reasons for abandonment
Recommendation
Set expectations for data quality and technical requirements (e.g. computing capabilities) on the Library.Harvard page
Insight 3
Finding specific files was very difficult even for expert users
"I haven't published anything yet because building the pipeline to extract the metadata to locate books is taking a long time”
Expert user
Finding
Filtering and locating specific books were challenging. Both expert and non-expert users found folder and file names vague
Recommendation
Orient users to the Corpus with documentation on the file structure of the Corpus with with basic details about folders
What this means
Users abandon the Corpus because they don't know what to expect, where to find the data they need, or how to get help
Impact
I pitched my recommendations to drive Corpus utilization for non-expert users, securing approval from the team and UX & Discovery Director
Both expert and non-expert users experience similar friction using the Corpus. Addressing the needs of non-expert users will streamline the workflow of expert users, driving utilization.
Set expectations
surface data quality and technical requirements for transparency
Orient users
document the Corpus's file structure so users can find what they're looking for
Make help easy to find
FAQs, video tutorials, or a persistent contact point for support
" Ally, something I've noticed over the course of the internship is that you are someone who often catch something, the rest of the team might not have considered and the projects comes out better because your questions. Your persistence to learn about the LLM project despite all it's challenges was very impressive. "
Meg McMahon, UX Manager
Takeaways
If I were to continue this project, I would keep recruiting with targeted recruitment with Harvard affiliates until I've reached enough participants (n=15)
Leverage AI to explain technical concepts that appear in interviews and surveys to gain subject matter understanding
Design a survey to still get insights in the case that we do not hear back from many participants
Bring in stakeholders early and often throughout the research process to validate direction