Research Projects

Project

Gigatoken Lab

Independent reproduction of Marcel Rød's Gigatoken tokenizer claims on Apple Silicon, with an educational recreation of the SIMD pretok and hierarchical cache levers.

Status Active

Source repository ↗

Gigatoken Lab

Do Gigatoken’s published speed claims — roughly 500–1000× vs HuggingFace tokenizers and ~100× vs tiktoken — survive on a different Apple Silicon machine, and can the two systems levers (SIMD pretok, hierarchical pretok cache) be recreated enough to explain why? Upstream marcelroed/gigatoken remains the production library; this lab owns reproduction benches, a hash-pinned results.json, and a Marimo notebook that demonstrates regex pretok vs a hand FSM, Zipfian cache capacity, and a live side-by-side encode.

On Apple M5 Max against a 200 MB OpenWebText slice, warm GPT-2 encode hits 4453 MB/s vs 33.4 MB/s HuggingFace — 133×, validated on 20,401 documents. Against tiktoken the subset path is 42× (4953 vs 118 MB/s) with zero token-ID mismatches in the first 50 documents; the native file API reaches 6.9 GB/s on the full 210 MB slice. Absolute ratios sit below the published M4 Max table (1,268× HF / 140× tiktoken); corpus size, cache residency, and API path move the number, but the order-of-magnitude claim vs HuggingFace holds.

Public surfaces are the pinned reproduction report, notebook snapshot, and artifact hash. Training corpora and machine-local bench trees stay private; the lab does not fork or replace upstream Gigatoken.

Reports

  1. Report

    Gigatoken Reproduction — SIMD Pretok and Hierarchical Cache

    Independent Apple M5 Max reproduction of Gigatoken's speed claims: 133× vs HuggingFace on GPT-2, 42× vs tiktoken, exact token-ID match — plus an educational recreation of the two systems levers.

    Jul 2026 Positive Notebook
  • tokenization
  • systems
  • simd
  • benchmarks