Skip to main navigation Skip to search Skip to main content

Large language model tokens are psychologically salient

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Large language models segment words into chunks called tokens, using compression algorithms that ignore semantics. We investigated whether tokenization corrupts representations of word meanings in 17 languages. We found that GPT-4o and Llama 3 inflate the similarity of words that share tokens. However, tokens turned out to be good predictors of orthographic priming, such that people recognize a target word faster after reading a prime that ends with the same token. This boost in priming far exceeds what other overlapping strings of letters explain, which suggests that tokenization selectively identifies functional subword units. The pattern extends to the production of word associates in English: Tokens capture phonologically motivated associations, while other strings of letters do not. So, tokenization does influence semantic representations, but because tokens correspond to psychologically salient orthographic and/or phonological constituents, they may endow large language models with human-like language networks and facilitate alignment with human word processing. ©2025 the author(s).
Original languageEnglish
Title of host publicationProceedings of the 47th Annual Conference of the Cognitive Science Society
EditorsD. Barner, N.R. Bramley, A. Ruggeri, C.M. Walker
PublisherUniversity of California
Pages4819-4827
Number of pages9
Publication statusPublished - Jul 2025
Event47th Annual Meeting of the Cognitive Science Society (CogSci 2025) - San Francisco, United States
Duration: 30 Jul 20252 Aug 2025

Publication series

NameProceedings of the Annual Meeting of the Cognitive Science Society
Volume47
ISSN (Electronic)1069-7977

Conference

Conference47th Annual Meeting of the Cognitive Science Society (CogSci 2025)
PlaceUnited States
CitySan Francisco
Period30/07/252/08/25

Research Keywords

  • large language models
  • tokenization
  • conceptual alignment
  • semantic priming
  • subword processing

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

Fingerprint

Dive into the research topics of 'Large language model tokens are psychologically salient'. Together they form a unique fingerprint.

Cite this