By 2020, language-powered features - search bars, tag clouds, voice assistants, auto-translate - were already built into almost every product we used daily. This was NLP, years before LLMs made AI mainstream. UX testing methods hadn't caught up: they were built mostly for visual interfaces, not for systems where the interaction itself is language.
I set out to answer two questions:
- How do people actually behave when interacting with linguistic interface components?
- Can established UX methods be adapted to evaluate these systems effectively?
This thesis bridges UX research, cognitive psychology, and NLP - building a methodology for testing systems where language is the primary interface.