30
UN Digital Library
Designing for What Users Actually See
The United Nations Digital Library hosts one of the world’s largest institutional archives — but its most-visited page was quietly failing the people trying to use it. Researchers struggled to interpret specialist terminology, layouts shifted unpredictably across record types, and the platform’s analytics were so bot-saturated that the UNDL couldn’t reliably measure real user behavior. Before we could evaluate the interface itself, we first had to rebuild a trustworthy picture of how people were actually using it.
I led the analytical layer of the study — cleaning and reconstructing behavioral data, operationalizing usability metrics, and translating raw session outputs into defensible findings.
- Two key findings consolidated from the 36-issue rainbow sheet — (1) specialist labels and terminology create access barriers; (2) inconsistent layouts and sub-WCAG typography compound them
- Four implementation-ready recommendations scoped against the existing vendor relationship — 'Formats' → 'Citation Formats,' a tooltip pattern for terms locked to internal taxonomy, a disabled-state Download pattern for non-digitized records, and a 14px / 400-weight typography update for WCAG compliance
- Final report delivered to the UNDL's Chief of Information Management Section, with a pending invitation to present findings to the broader library and tech department
Data reveals where users struggle. The real work is making the reason impossible to ignore.
Every heatmap gap is a specific, diagnosable failure — a visual hierarchy problem, a contrast issue, an affordance that doesn't communicate what it needs to. But that diagnosis only holds if the underlying data is clean, consistently structured, and contextualized against something real. The comparative benchmark didn't just add a slide to the presentation — it changed how stakeholders heard the SUS score. Numbers land differently when they're anchored. That's the work.
A successful task and a confident user aren't the same thing.
Five of eight participants completed the linked-record task correctly. Every one of them, including the five who succeeded, rated the task at 1.9 / 7 difficulty and reported they didn't trust they'd found the right thing. Task completion is a binary the interface measures easily. Confidence is the variable it actually changes. When eye-tracking and self-report tell different stories about the same moment, that gap is where the design work lives.
Specialist language isn't neutral — it's an access barrier with traffic numbers attached.
The UNDL's labels weren't wrong; they were internally consistent with a bibliographic system built by experts for experts. But every general-public user who left without finding a document was a measurable cost of that consistency. The recommendation — relabel where you can, tooltip where you can't — is small. The principle is bigger. Every specialist label assumes the user already speaks the system’s language.