[Demo] How to re-categorize content at scale using LLMs
Summary
Large Language Models (LLMs) are to language as spreadsheets are to numbers: tools for modeling, exploration, and development. Among their many capabilities, LLMs can alleviate chores related to the design and implementation of information architectures. But doing so requires venturing beyond chat-based interfaces. In this brief demonstration, we'll see how to use OpenAI's API and a few open source command line tools to re-categorize content in a 1,000+ page website. The techniques demonstrated can be extended to other common content organization tasks.
Key Insights
-
•
Manual retagging of 1,200 blog posts would take about 10 hours, but leveraging GPT-4 reduced active human time to about 2 hours.
-
•
Using GPT-4 via command line and shell scripts enables automated tagging outside typical chat interfaces.
-
•
An organically grown taxonomy over 20 years contained unclear acronyms and inconsistent tag forms that GPT initially struggled with.
-
•
Cleaning and standardizing the taxonomy before prompting GPT is critical for effective AI assistance.
-
•
A review step of AI-suggested tags in CSV format allows human correction to avoid hallucinations entering production.
-
•
GPT-4 can propose new and useful tags outside the original taxonomy, enriching content classification.
-
•
The four-step GRU framework (Gather, Review, Update, Wrap up) balances automation with human oversight.
-
•
Storing blog content as markdown files simplifies integrating AI workflows via scripting and file manipulation.
-
•
The approach is adaptable and scalable to other CMS platforms by replacing scripting with API calls.
-
•
Taxonomies should use clear, unambiguous terms to improve both human and AI understanding.
Notable Quotes
"Some of the older content has discoverability problems, which is typical with blogs."
"Doing this tagging manually would have taken me around 10 hours of mind-numbing work."
"I’m actually using GPT-4, but not via the chat interface—I'm calling it from the Mac’s command line."
"I had to clean the taxonomy up because GPT wouldn’t know what to do with acronyms like TAOI."
"I save the proposed tags to a CSV file so I can preview and edit them before applying the changes."
"A middle review step prevents hallucinations from making it into the production site."
"GPT-4 functioned as an assistant not just in retagging but also in improving the taxonomy itself."
"The entire process took about three hours from start to finish, about a fifth of the manual time."
"Use clear and obvious terms in taxonomies—unusual acronyms won’t make sense to GPT or others."
"You need to review proposed changes before committing them to production, otherwise errors sneak in."
Or choose a question:
More Videos
"Good design needs better PR and the government has to be actively involved in the process."
Sofía Delsordo Kassim VeraPublic Policy for Jalisco's Designers to Make Design Matter
December 8, 2021
"We are like three legs of a chair: product, engineer, and UX all contribute to stability."
Product and Design at Bloomberg: A 15-year Evolution
December 6, 2022
"Research is a sales role at its heart, delivering value to people to drive sustainable growth."
Liza Pemstein Jane DavisScaling Research Via an Ops First Model at Clever
March 27, 2023
"High-level leadership changes in government often result in shifts in organizational moods and priorities."
Kara KaneTheme One Intro
November 16, 2022
"The risk of AI is that it will over-rely on past priorities and miss new, emerging signals."
Mark Interrante Harry MaxAI for Prioritization (3rd of 3 seminars)
July 11, 2024
"Confidence is critical in online shopping because nobody wants to spend money they’re unsure about losing due to unclear controls."
Sam ProulxOnline Shopping: Designing an Accessible Experience
October 3, 2023
"We try to figure out who needs to be in the room to hear it — research is how we build team cohesion."
Rich MironovHow Can Product Managers and UXers Help Each Other (and Why are Product Folks so Annoying Sometimes)?
December 6, 2022
"For every dollar spent on cancer care, roughly four are spent on administration."
Juhan SoninDesign Now! The Agenda for Action
September 4, 2025
"The trust gap stems from being repeatedly failed by a system, not innate risk aversion or imposter syndrome."
Mansi GuptaDrawing from Feminist Practice to Make Inclusive Design Operational
September 9, 2022