搜索 — ResearchTracker

Attention mechanisms are central to the success of large language models (LLMs), enabling them to capture intricate token dependencies and implicitly assign importance to each token. Recent studies have revealed the sink token, which receives disproportionately high attention despite their limited semantic role. In this paper, we first expand the relationship between the sink token and other tokens, moving beyond attention to explore their similarity in hidden states, considering the layer depth. We observe that as the layers get deeper, the cosine similarity between the normalized hidden states of the sink token and those of other tokens increases, and that the normalized hidden states of the sink token exhibit negligible changes. These imply that other tokens consistently are directed toward the sink token throughout the layers. Next, we propose a dynamic token selection method, called OrthoRank, using these findings to select important tokens. Specifically, in a certain layer, we define token importance by the speed at which the token moves toward the sink token. This is converted into orthogonality with the sink token, meaning that tokens that are more orthogonal to the sink to

Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece

arXiv2026-01-05作者：Anshul Kumar

Tokens are the basic units of Large Language Models (LLMs). LLMs rely on tokenizers to segment text into these tokens, and tokenization is the primary determinant of computational and inference cost. Sanskrit, one of the oldest languages, is hypothesized to express more meaning per token due to its morphology and grammar rules; however, no prior work has quantified this. We use a dataset of 701 parallel verses of the Bhagavad Gita, which comprises three languages-Sanskrit, English, and Hindi along with transliteration of Sanskrit into English. We test tokenizers including SentencePiece (SPM), older GPT models, and the latest generation tokenizers from Gemini and GPT. We use metrics of token count, characters per token (token efficiency), and tokens per character (token cost). Results show a ~2x difference in token counts between Sanskrit and English/Hindi under the unbiased SPM baseline. English/Hindi translations of Sanskrit commentary resulted in an approximately 20x increase in token count. GPT o200k base (latest, used by GPT-4o) and Gemini (latest) reduce bias by a significant degree compared to GPT cl100k base (used until GPT-4), but still fail to fully capture Sanskrit's comp

搜索结果：Token

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece

Explaining and Mitigating Crosslingual Tokenizer Inequities

VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models

Context-Aware Wireless Token Communication via Joint Token Masking and Detection

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

Rethinking Tokenized Graph Transformers for Node Classification

Exploring Token-Space Manipulation in Latent Audio Tokenizers

Recursive Augmented Fernet (RAF) Token: Alleviating the Pain of Stolen Tokens

Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing

ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

Not all tokens are needed(NAT): token efficient reinforcement learning

Tokens with Meaning: A Hybrid Tokenization Approach for Turkish

Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers

On Generalized Token Graphs

Beyond Literal Token Overlap: Token Alignability for Multilinguality

Token Weighting for Long-Range Language Modeling

Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production