Abstract:
A token is the discrete unit for data processing in recent natural language processing techniques. With the rapid development of large language models, tokens provide a unified representation for diverse modalities, enabling cross-modal understanding and generation. From text subwords to visual patches, tokenization enhances data processing efficiency but also faces challenges such as semantic alignment and information loss. This article introduces the background, overview, and future directions of tokenization technology.