Researchers have developed Joint Optimization for Greedy Longest-Match Tokenization (JOLT), a new algorithm that outperforms the widely used Byte Pair Encoding (BPE) method for breaking text into tokens that AI models can process.

The work, published by a team led by Adhiraj Singh and Deepanshu Mody, addresses a fundamental bottleneck in how language models handle text. Current tokenization methods like BPE rely on greedy heuristics rather than optimizing for the specific way models decode tokens during inference.

JOLT formulates vocabulary learning as an integer programming problem, using mathematical constraints to ensure the training process matches exactly how tokens are processed during deployment. The approach specifically targets WordPiece's greedy left-to-right longest-match decoding, which powers many production AI systems.

Performance gains across model scales

Testing across vocabulary sizes of 32,000 and 64,000 tokens, JOLT produced up to 0.78% fewer tokens than BPE on validation data. The improvements grew larger as training scope increased, suggesting the method scales effectively with model size.

The researchers found that BPE already achieves within 1-2% of optimal compression under greedy longest-match decoding. However, JOLT closes 89.6-99.4% of the remaining gap between BPE and theoretical optimality.

To make the optimization computationally tractable, the team developed a linear programming relaxation that selectively introduces complex segmentations only where needed. The resulting solutions fell within 0.008-0.176% of the mathematical lower bound.

The work provides both practical improvements and theoretical insights into tokenization limits. By aligning training objectives with deployment-time behavior, JOLT demonstrates that inference-aware optimization can extract meaningful gains even from mature techniques like BPE.

The research comes as tokenization efficiency becomes increasingly important for AI companies managing inference costs at scale. Better compression directly translates to reduced computational requirements and faster model responses.