Speck is a compact research harness for experimenting with small causal language model architectures, training them, benchmarking optimization performance, and running checkpoint inference.
Models are groups of residual blocks. Each block contains ordered stages, and a stage can run one or more branches in parallel. Supported architecture components include:
- Global or sliding grouped-query attention.
- Gated causal convolution.
- SwiGLU feed-forward layers.
- Repeated blocks with optional weight sharing.
- Heterogeneous block widths and attention head dimensions.
Speck requires Python 3.10 or later and uses uv. Create either a GPU or CPU environment:
uv sync --extra gpu
# or
uv sync --extra cpuActivation is optional when using uv run. For an interactive shell, use the command appropriate to your shell:
# POSIX shells
source .venv/bin/activate
# fish
source .venv/bin/activate.fishChecked-in experiment directories identify an architecture and its data, tokenizer, and training configuration:
experiments/Speck1-140Mis the 140,652,288-parameter production and architecture-search configuration.experiments/Speck1.5-140Mkeeps that architecture, tokenizer, and 5B-token optimization recipe while using its isolated stationary corpus mixture.experiments/Speck1-140M-Instructreuses that base architecture and tokenizer to post-trainSpeck1-140M-Instructon SpeckChat1.experiments/Speck1.1-140M-Instructreuses that base architecture and tokenizer to post-trainSpeck1.1-140M-Instructon SpeckChat2 for one epoch.experiments/Speck1.1-140M-Instruct-2epretains the corresponding two-epoch training run.
Model names follow Speck<generation>-<size>, with an optional decimal generation for intermediate families.
A search-capable experiment directory contains five JSON files:
model.json Architecture and dimensions.
tokenizer.json Tokenizer artifact and local directory.
data.json Sources, phased mixture, filters, dedup, shards, and packed output.
train.json Optimization, batching, logging, and checkpoints.
search.json Search space, training rungs, scoring, and profiling contract.
Artifacts use ~/.cache/speck by default. Checkpoints are written to ~/.cache/speck/checkpoints/<train.run>. Packed data uses an explicit data.output_dir, ~/.cache/speck/data/<data.output_name> when a name is configured, or the legacy ~/.cache/speck/data/packed default. Set speck_base_dir to move the cache root.
Despite its historical name, train.json's min_lr is a multiplier of the peak lr, not an absolute learning rate. For example, 0.1 ends the schedule at 10% of the peak rate.
Download and verify the tokenizer configured by an experiment:
python -m scripts.tokenizer_prepare experiments/Speck1-140MResolve, stream, filter, deduplicate, tokenize, and pack the configured sources:
python -m scripts.data_prepare experiments/Speck1-140MPrepare the isolated Speck1.5 corpus under ~/.cache/speck/data/Speck1.5-140M-corpus:
python -m scripts.data_prepare experiments/Speck1.5-140MThe base pretraining experiment requests 5,000,000,000 training tokens. Its phase schedule is:
| Phase end | ultra_fineweb | dclm | cosmopedia_v2 | finemath_4plus | ultrafineweb_l3 |
|---|---|---|---|---|---|
| 3,500,000,000 | 45% | 35% | 12% | 8% | 0% |
| 4,500,000,000 | 30% | 25% | 15% | 12% | 18% |
| 5,000,000,000 | 20% | 15% | 20% | 15% | 30% |
The phase durations and integer weights derive source targets of 1.975B, 1.55B, 670M, 475M, and 330M tokens respectively. Preparation adds a derived 262,144-token per-source loader reserve for the configured maximum 65,536-token distributed microbatch, then reports each requested target, reserve, and actual full-document result. Actual packed training data can exceed 5B only by these configured reserves and one final full-document overshoot per source.
The Speck1.5 corpus uses one stationary mixture for exactly 5B requested training tokens:
| Source | Tokens | Share |
|---|---|---|
| FineWeb-Edu | 2.500B | 50% |
| DCLM-Edu | 1.650B | 33% |
| FineMath-4+ | 350M | 7% |
| UltraData-Math L3 Textbook-Exercise | 75M | 1.5% |
| UltraData-Math L3 Multi-Style | 25M | 0.5% |
| Wikimedia | 125M | 2.5% |
| peS2o | 175M | 3.5% |
| Ultra-FineWeb-L3 Multi-Style | 50M | 1% |
| Cosmopedia v2 | 50M | 1% |
The category totals are 83% natural web, 9% math, 6% knowledge/science, and 2% general synthetic. DCLM-Edu retains English rows with strict raw edu_score > 3.5; the two mixed-language UltraData-Math configurations retain rows identified as English by pinned py3langid==0.3.0. All repository revisions and dataset paths are pinned in data.json.
The previous corpus and stopped run are obsolete and are not resumed or reused. The current corpus uses distinct packed-data and checkpoint names.
Repository revisions are resolved once and pinned, and recursive Parquet discovery uses the Hugging Face repository tree rather than datasets-server previews. Files are deterministically shuffled per source. Preparation downloads and reads only one remote Parquet file at a time, removes it immediately, and writes train and validation shards under sources/<source-id>/. Validation reserves 5M tokens per source and the loader schedules those streams equally.
Exact global deduplication normalizes text with Unicode NFKC, lowercasing, and whitespace collapse before recording a 128-bit BLAKE2 hash. The expected roughly 6M hashes remain practical in memory and are journaled compactly at 16 bytes each. A collision is treated as a duplicate; fuzzy and LSH deduplication are intentionally excluded. Tokenizer calls are bounded to 1,024 documents and 2,000,000 aggregate input characters.
Preparation performs a live disk-space preflight before creating staged data. The current estimate includes about 10.05GB of packed uint16 data, a 20GiB temporary raw-shard allowance, and at least 5GiB of dedup/index headroom, for about 36.9GB total required capacity. The command reports required and currently free bytes and credits reusable staged bytes on resume.
Preparation builds under the sibling .building directory and atomically publishes the final directory. Every completed remote Parquet file closes and checkpoints packed shards, source-local index bytes, and the dedup journal with checksums. A retry validates those boundaries, removes only partial work from the interrupted file, and resumes at the next file. Pass --restart to discard all staged state. A completed output directory is never overwritten.
Authenticate with Weights & Biases, then start a single-GPU run:
wandb login
python -m scripts.base_train experiments/Speck1-140MUse experiments/Speck1.5-140M to train against its isolated packed corpus with the same command structure.
Weights & Biases logging is enabled unless train.run is dummy. Checkpoints remain local; Speck does not upload training checkpoints to Hugging Face.
Launch distributed data-parallel training with torchrun:
torchrun --standalone --nproc_per_node=8 -m scripts.base_train -- \
experiments/Speck1-140MThe configured 65,536-token optimizer batch is divisible by device_batch_size * sequence_length * world_size for world sizes 1, 2, 4, and 8. Since 5B is not batch-aligned, training performs 76,294 optimizer steps and consumes 5,000,003,584 tokens. Mixture phases are selected from each global microbatch's starting token position, so a microbatch that begins before a phase boundary remains in that phase even if it straddles the boundary.
Existing checkpoints are never resumed implicitly. A run fails rather than overwrite them unless an exact checkpoint step is supplied:
python -m scripts.base_train experiments/Speck1-140M --resume <checkpoint-step>Resume validates the architecture, packed-data manifest, optimizer settings, batch geometry, training horizon, world size, and that the next-batch loader offset exactly equals completed optimizer-step tokens. It restores the optimizer, data position, elapsed time, and W&B run identity.
Prepare the pinned specklabs/SpeckChat1 dataset with the Speck chat template and assistant-only loss mask:
python -m scripts.sft_prepare experiments/Speck1-140M-InstructThe current SpeckChat1 post-training configuration uses <|system|>, <|user|>, and <|assistant|> as token IDs 32000-32002, preserves the pretrained BOS/EOS tokens, and holds out 1,000 conversations for validation. Conversations are isolated in 256-, 512-, 1,024-, or 2,048-token buckets. The per-device batches are 32, 16, 8, and 4 respectively, so every microbatch has the same 8,192-token compute budget without unnecessary 2,048-token padding. Start one epoch of full-model instruction tuning from the pinned specklabs/Speck1-140M release:
python -m scripts.sft_train experiments/Speck1-140M-InstructUse torchrun as with base training for multiple GPUs. SFT checkpoints and a Hugging Face-compatible tokenizer are written under ~/.cache/speck/checkpoints/Speck1-140M-Instruct. Resume only from an explicit SFT step with --resume <checkpoint-step>.
Generate from the instruction-tuned checkpoint by selecting its directory. The prompt is automatically rendered as a user message when the checkpoint metadata identifies SFT:
python -m scripts.infer "Explain why the sky is blue." \
--checkpoint-dir ~/.cache/speck/checkpoints/Speck1-140M-InstructRebuild and publish the 500,000-row specklabs/SpeckChat2 train split with pinned source revisions, source-specific quality filters, exact prompt deduplication, and Speck-tokenizer length checks:
uv run scripts/speckchat2_prepare.pyThe mixture contains 200K LMSYS DeepSeek conversations, 130K Magpie Llama 3.1 multi-turn conversations, 85K Hermes, 65K UltraChat, 10K Magpie Reasoning, 8K No Robots, and 2K Everyday Conversations. It uses only source training splits and intentionally publishes no validation or test split. Use --output-dir <path> --no-push to build a local dataset instead.
The experiments/Speck1.1-140M-Instruct configuration pins the published 500,000-row SpeckChat2 dataset and the original Speck1-140M base weights. Prepare its isolated assistant-masked data, holding out 1,000 conversations for validation:
uv run --extra gpu python -m scripts.sft_prepare experiments/Speck1.1-140M-InstructRun one epoch of full-model post-training:
uv run --extra gpu python -m scripts.sft_train experiments/Speck1.1-140M-InstructPrepared data is written under ~/.cache/speck/data/SpeckChat2-v3, and checkpoints are written under ~/.cache/speck/checkpoints/Speck1.1-140M-Instruct.
Run the retained two-epoch variant against the same prepared data:
uv run --extra gpu python -m scripts.sft_train experiments/Speck1.1-140M-Instruct-2epIts checkpoints are written under ~/.cache/speck/checkpoints/Speck1.1-140M-Instruct-2ep.
Generate from the latest checkpoint, or select one with --step:
python -m scripts.infer "The meaning of life is" \
--experiment experiments/Speck1-140MUseful controls include --max-tokens, --temperature, --top-k, --device, and --checkpoint-dir.
Export, validate, and publish the canonical one-epoch instruction checkpoint as a BF16 Transformers repository:
uv run --extra cpu python -m scripts.model_publish --expected-epochs 1Published likelihood evaluation supports binary right-padded batches when use_cache=False.
Left padding, mask gaps, and cached padded inference remain unsupported. Apply the same tracked
compatibility code to the existing base-model repository without changing its weights:
uv run --extra cpu --with transformers==5.1.0 python -m scripts.model_code_publishThe code-only publisher verifies the immutable source, model-weight LFS checksum, Auto class
loading, parameter count, padded-batch logit parity, uploaded code hashes, and unchanged remote
weights. Use --no-upload to run every local validation without creating a Hub commit.
Build BF16, Q4_K_M, Q5_K_M, and Q8_0 GGUF variants of the published instruction model, smoke-test each file with a pinned llama.cpp checkout, and publish them to the Hugging Face Hub:
uv run --extra cpu python -m scripts.gguf_publishGenerated weights and the llama.cpp checkout are kept under ~/.cache/speck, not in this
repository. Use --no-upload for a local build, repeat --quantization <type> to select a
different set, or pass --llama-cpp <path> to use an existing checkout. The publisher resolves
the source model to an immutable revision and uploads only after every requested file passes a
llama.cpp load and inference smoke test. Work is capped at four concurrent jobs by default; use
--resume after an interruption to validate and reuse completed files.
Run every benchmark in the pinned Open SLM Leaderboard configuration against the immutable
specklabs/Speck1-140M release:
uv run --extra gpu --group open-slm python -m scripts.open_slm_eval allThe configuration in experiments/Speck1-140M/open_slm.json pins the leaderboard, model,
lm-eval harness, standard-task datasets, and both official ArithMark repositories and file
checksums. Results default to ~/.cache/speck/evaluations/open-slm/Speck1-140M. Use the
individual lm-eval, arithmark-2, arithmark-3, and summary stages to resume a run, or
pass --limit 2 to lm-eval for a smoke test. ArithMark 2.0's verified official runner
right-pads without disabling the model cache; the wrapper leaves its scoring code unchanged
and sets model.config.use_cache=False immediately after model loading.
Pinned zero-shot results are recorded under results/<model>/open_slm.json:
| Model | HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 | Int Index | ArithMark-2 |
|---|---|---|---|---|---|---|---|
| Speck1-140M | 35.03 | 46.68 | 25.94 | 63.87 | 36.60 | 18.15 | 31.52 |
| Speck1-140M-Instruct | 35.22 | 45.66 | 25.85 | 63.60 | 36.10 | 17.75 | 33.64 |
| Speck1.1-140M-Instruct | 35.64 | 46.93 | 26.02 | 64.15 | 33.70 | 17.90 | 32.44 |
Evaluate both public instruct releases with the same raw-continuation tasks and no chat template:
uv run --extra gpu --group open-slm python -m scripts.open_slm_eval all \
--config experiments/Speck1-140M-Instruct/open_slm.json
uv run --extra gpu --group open-slm python -m scripts.open_slm_eval all \
--config experiments/Speck1.1-140M-Instruct/open_slm.jsonThe model-specific configs inherit the benchmark and dataset pins from the base evaluation config. Their outputs default to separate directories named after each Hub repository.
Measure compiled optimization steps with synthetic input:
python -m scripts.benchmark experiments/Speck1-140M \
--mode compute \
--output benchmark.jsonUse --mode end-to-end --data-dir ~/.cache/speck/data/packed to include packed-data loading. Warmup is reported separately. --peak-tflops reports model FLOPs utilization, and --no-compile measures eager execution.
Run the gated BananaMind Base Bench 1.1 continuation benchmark with its pinned official runner:
python -m scripts.bananamind_bench \
--model experiments/Speck1-140M \
--speck-checkpoint-step 76294 \
--device cuda \
--dtype bfloat16 \
--batch-size 32Accept the dataset gate and authenticate with Hugging Face first. The wrapper verifies the runner and data checksums, pins checkpoint and tokenizer hashes in the report, and rejects resume when the checkpoint or numerical configuration changes. It delegates scoring to the official runner and requires transformers in the execution environment.
Compare normalized prompt-prefill and cached-decoding inference speed with pinned model revisions:
python -m scripts.inference_benchmark --model speck --device cpu
python -m scripts.inference_benchmark --model speck --device cudaSelect speck, supra, gptx, banana, or smol. CPU defaults to FP32 batch 1; CUDA defaults to BF16 batches 1 and 32. The benchmark excludes tokenization, uses eager SDPA, returns only the final-position logit, and records raw synchronized timings. External models require transformers.
uv run --extra cpu --group dev pytest -qSearch uses the baseline experiment and stores a resumable study under ~/.cache/speck/search/<name>:
python -m scripts.search run experiments/Speck1-140M \
--name evolution-01 \
--hours 3
python -m scripts.search status evolution-01
python -m scripts.search finalize evolution-01See Architecture search for prerequisites, runtime contracts, promotion behavior, output files, and finalization.