Complete AI Training

AI news ·

Researchers test fixed token codes at 100 billion token scale, find they work without trainable embeddings

Fixed token codes rival learned embeddings in three 1.7B-parameter language models trained on 100B tokens. One fixed-code variant hit 52.40% on HellaSwag and cut trainable parameters by 100.7 million.

Share

A team of researchers trained three 1.7-billion-parameter language models from scratch on 100 billion tokens, finding that fixed token codes can perform competitively with learned embeddings. The results, published October 6 on arXiv, challenge the assumption that trainable input vectors are architecturally necessary for decoder-only models to function well.

The three models shared identical tokenizers, Transformer backbones, and training recipes. They differed only in how they converted tokens into vectors fed to the first layer. One used a standard learned embedding table. The second used canonical 16-bit token-ID codes. The third employed a fixed invertible recoding over GF(2).

Performance held up across standard benchmarks

The fixed-code models delivered solid results. One fixed-code variant reached 52.40% on HellaSwag, 70.51% on PIQA, and 42.75% on LAMBADA. The learned-embedding model outperformed the fixed versions on some tasks, but the gap was not large enough to dismiss the fixed approach.

The fixed interfaces cut trainable parameters by 100.7 million compared to the learned embedding baseline. The paper emphasizes this parameter reduction is not the main finding. The core question was whether token-specific input vectors are essential or merely an empirical convenience.

What the study set out to test

The work aims to clarify a design assumption baked into nearly every large language model. Learned embeddings map each token in the vocabulary to a dense vector that the model adjusts during training. Fixed codes bypass that mapping entirely, replacing it with a deterministic, non-learned representation.

The authors trained all models from scratch at scale - 100 billion tokens - to ensure differences in performance reflected architectural choices rather than insufficient training budgets. The strong benchmark scores from fixed-code models suggest that much of what embeddings contribute can be recovered by the model's later layers.

Why this matters for R&D engineers and developers

For teams building or fine-tuning language models, the findings point toward simpler input pipelines. Removing the embedding table eliminates a large parameter block and the associated memory and initialization complexity. That matters when you are optimizing for inference speed, deployment footprint, or training on constrained hardware.

The paper does not claim fixed codes beat learned embeddings. It shows they are viable. Engineers working on domain-specific models or edge deployments may find the trade-off useful - competitive performance with fewer moving parts. For those looking to deepen their understanding of model internals, Generative AI Courses cover Transformer architectures and embedding design choices in detail. Professionals building research-track skills can also explore the AI R&D Engineering Courses learning path.

Share