User Research on AI Search

A master's thesis on how users navigate and search AI-powered interfaces, back in 2020 before AI was everywhere · 7 min read
Research Question
By 2020, NLP-powered components were already built into almost every product we used, but standard UX methods were still mostly designed for visual interfaces, not for systems whose core interaction happens through language. How do people actually behave when interacting with linguistic interface components back then?
Context
Master's thesis · Adam Mickiewicz University · 2020
Role
Sole Researcher - designed and ran the full study: survey creation, participant recruitment, usability test facilitation, Wizard of Oz sessions, and data analysis.
UX Research Usability Testing Survey Design Data Analysis
Outcome
5 actionable findings and 4 design principles for linguistic interface UX, validated across 3 live systems with 88 survey respondents and 33 usability-test participants.
Duration
6 months
User Research on AI Search

Research Question

By 2020, language-powered features - search bars, tag clouds, voice assistants, auto-translate - were already built into almost every product we used daily. This was NLP, years before LLMs made AI mainstream. UX testing methods hadn't caught up: they were built mostly for visual interfaces, not for systems where the interaction itself is language.

I set out to answer two questions:
  • How do people actually behave when interacting with linguistic interface components?
  • Can established UX methods be adapted to evaluate these systems effectively?


This thesis bridges UX research, cognitive psychology, and NLP - building a methodology for testing systems where language is the primary interface.

By the Numbers

88
survey
respondents
3
systems tested
33
usability test
participants

Participant Mix

The 88 survey respondents were an international group: about half in Poland, the rest across six other countries, mostly aged 21–40. This split underpins the later Poland-vs-international comparisons in translation trust and voice-assistant use.
  • Poland 47.7%
  • United States 23.9%
  • Ukraine 18.2%
  • Macedonia 3.4%
  • Brazil 2.3%
  • Portugal 2.3%
  • Italy 2.3%
  • 21–30 55.7%
  • 31–40 25.0%
  • 41–50 14.8%
  • Under 20 4.5%
Survey respondents by location and age

Research Process

A two-phase mixed-methods study grounded in ISO 9241 human-centred design principles, cognitive psychology, and social psychology.
01
Large-Scale Survey
88 respondents across an international group, exploring everyday habits with search, hashtags, machine translation, and voice assistants.
02
In-Depth Interviews
Selected participants for 30-minute deep interviews to uncover motivations behind their survey answers.
03
Usability Testing
Task-based tests on three real systems: three rounds with limited search options on the main system, plus one open-ended round on each of the two other systems.
04
Wizard of Oz Sessions
I stepped in as a hidden assistant whenever a user got stuck, revealing features they'd missed. Each round followed the same loop: watch and learn, build, test with real users, step in when something broke, repeat.
Start Watch & learn Build Test live Test live, with backup assistant Repeat
Extended Wizard of Oz model adapted for linguistic interface testing
Small test groups were a deliberate choice. Nielsen and Landauer's classic curve shows that five users typically surface 85% of a system's usability problems - beyond that, sessions mostly re-confirm what you already know.
0 25 50 75 100 0 5 10 15 Number of test participants Problems found (%) 5 testers → 85%
Usability problems found rise quickly, then flatten out around 5 testers

Systems Under Study

I selected three real-world systems that each represent a different approach to linguistic interface design - from text-based semantic search to visual tag navigation.
Pasaż Wiedzy Wilanów
A knowledge portal about Polish Baroque culture featuring a semantic search engine, a thematic tag cloud, category-based navigation [1], and article-level text hashtags. I ran three rounds of testing here with user groups under different search constraints.
IKEA
Selected for its category navigation system [2]. Product categories are presented as hashtags, giving users a visual browsing path as an alternative to the traditional search bar.
Fragrantica
A fragrance knowledge portal with rich graphical tag filtering [3] - scent notes displayed as visual icons, plus an advanced search that lets users combine notes to include and exclude at once. Chosen to compare graphical hashtags against the text-only hashtags tested on Pasaż Wiedzy.
System Search Bar Text Tags Visual Tags Tag Cloud Auto-suggest
Pasaż Wiedzy
IKEA
Fragrantica
Scroll to explore
Pasaż Wiedzy Wilanów - tag cloud, search bar, and category navigation
Pasaż Wiedzy - text hashtags, tag cloud, and semantic search
IKEA category dropdown - text-based subcategory navigation
Text-link category menu for subcategory navigation on IKEA
Fragrantica graphical scent note tags
Graphical scent-note tags on Fragrantica

Test Design: Controlled Search Conditions

On Pasaż Wiedzy Wilanów, I divided participants into test groups with progressively restricted access to the search bar - forcing users to explore tag clouds, text hashtags, and category navigation they would normally ignore. Sessions averaged ~20 minutes across all four groups.
Group 1
Limited Search
One search allowed (single keyword or multi-word phrase). Users adapted quickly, relying on suggested articles and the sidebar tag cloud.
Group 2
No Search Bar
No search bar access at all. Sessions averaged ~30 min. Two participants abandoned the task entirely; others reported high frustration.
Group 3
Unlimited Search, No Suggestions
Full search access but the "related articles" list was disabled. Users only discovered text hashtags under articles after several minutes.
Group 4
Wizard of Oz Extension
During post-test interviews I revealed hidden features as a "system assistant" - most commonly the auto-suggest and tag-cloud sidebar.
Text hashtags under articles - easily overlooked by users
Suggested articles sidebar - users' preferred navigation tool
Pasaż Wiedzy's text hashtags and suggested-articles sidebar

Key Findings

Search habits are deeply ingrained
Among the 88 survey respondents, most defaulted to typing a single keyword into a search bar - Google alone accounted for over 90% of searches.
0 10 20 30 40 50 Google 41 47 Bing 2 0 DuckDuckGo 2 5 Yahoo 0 1 Yandex 0 2
  • Poland
  • International
Search engine choice, by region
68% never used text-based hashtags, and even when forced to try other tools, people raced back to keyword search.
  • No 68.2%
  • Yes 29.5%
  • Don’t know 2.3%
Text-hashtag usage among respondents
Visual tags outperform text tags
Users on IKEA and Fragrantica engaged with graphical hashtags significantly more readily than with the text hashtags on Pasaż Wiedzy. Graphical tags felt like "browsing" rather than "searching" - a mental model users were more comfortable with.
IKEA visual product carousel with tag-based tabs
Fragrantica advanced visual tag search
Graphical tags on IKEA and Fragrantica
Tag clouds are powerful but unintuitive
In interviews, users acknowledged that the thematic tag cloud expanded their search results meaningfully. However, they described it as "not visible at first glance" and "complicated for multi-tag queries." Better visual presentation was the top request.
Voice assistants: trusted for small tasks, not high-stakes ones
62% of respondents used voice assistants, rating satisfaction at 4/5. But in interviews, none said they'd trust one with financial tasks like paying bills or booking flights - fear of getting it wrong was the main barrier. The same caution shows up today with AI assistants generally: people delegate small things, but keep a human in the loop for anything that matters.
  • Yes 62.5%
  • No 37.5%
Voice assistant usage among respondents
Machine translation: widely used, moderately trusted
~70% of respondents used translation services, with Google Translate dominating. Average satisfaction was only 3/5 - heavy use, but a real trust gap.
0 5 10 15 20 25 1 1 9 2 2 10 24 3 18 12 4 1 2 5 Satisfaction rating (1 = low, 5 = high)
  • Poland
  • International
Translation satisfaction ratings, by region

Takeaways

NLP features rarely fail because the technology is bad - they fail because the interface doesn't make them discoverable or trustworthy. From the findings, four takeaways for designing systems with linguistic competence modules - built for 2020's search bars and tag clouds, but holding up just as well for today's AI chat interfaces:
01
Make NLP features visible
Tag clouds, auto-suggest, and related-content panels need to be front and center. If people don't spot a feature in the first few seconds, it may as well not exist.
02
Prefer visual over text for tags
Graphical representations of categories and tags lowered the barrier to use. Design tags as browsable visual elements, not text-only links.
03
Build trust incrementally
Voice assistants and translation tools are used daily but not fully trusted. Show provenance, confidence levels, or "why this result" to build user confidence.
04
Adapt UX methods for language
Standard usability testing misses linguistic interactions. Combine it with Wizard of Oz, controlled search constraints, and post-test interviews.

Reflection

What I Learned
I ran this study in 2020, back when AI still mostly meant recommendation engines and autocomplete, not something you could hold a conversation with. For a master's thesis, it was an ambitious bet: a full mixed-methods study run solo across three live systems, 88 survey respondents, 33 usability-test participants, and Wizard of Oz sessions, all in six months.

It's aged well. Six years later, the interfaces look nothing alike, but the behavior does. The same keyword-first habit I found in 2020 (68% of respondents ignoring text hashtags) is exactly what shows up now in how people prompt LLMs: short, search-style queries instead of exploring what the system can do.

What I Would Do Differently
I'd record sessions on video to review interaction patterns more rigorously. It's the closest substitute for eye-tracking, which wasn't in the budget for a university thesis.

Explore more projects

Investment Hub
Making mutual fund investing simple for retail investors.
Research UX Design Prototyping Dev Handoff
Read case study
Investment Hub
Sailes Charter Identity
Building a digital home for a family-run sailing charter company.
Research Brand Identity UI Design
Read case study
Sailes Charter Identity
Tile & Card Guidelines
Defining UX principles and usage guidelines for tiles and cards within a design system.
UI Design Design System
Read case study
Tile & Card Guidelines