Sol
Applied ML skills portfolio piece — train, measure, ablate, and ship a language model from scratch
- Role
- ML engineer
- Stack
- PyTorchTinyStoriesBPE TokenizationStreamlitHugging Face HubCUDAYAML Configs

Sol's demo sits behind a login — the app generates text on every request, and an open URL would let anyone drain its budget. Get in touch and I'll send you credentials. Request credentials
Why I built this
Sol is not a product trying to solve a user problem — it is a skills demonstration for applied ML / ML engineering roles. Recruiters and hiring managers evaluating me need evidence that I understand training dynamics, tokenization, evaluation design, and hardware trade-offs — not just how to call an API.
Fine-tuning a pretrained model is faster but weaker signal for fundamentals. Sol is the opposite: a small transformer trained properly and measured honestly, with every number guarded by tests and artifacts, so you can see how I work rather than what I shipped.
What it is
A 52,901,712-parameter decoder-only transformer, written by hand in src/model.py (~300 lines, nanoGPT-informed but not forked) and trained from scratch on TinyStories on a single RTX 4070 Laptop GPU (8 GB, BF16). No pretrained weights. The repository is the résumé — it contains the full stack:
- Data — download, dedup funnel, cross-split leakage detection, in-domain 32k byte-level BPE, binary shards, EDA notebook, data card
- Training — cosine LR schedule, grad clip, checkpoint/resume with RNG state, M4 smoke gates (overfit-one-batch, VRAM benchmark)
- Evaluation — full-val perplexity with 10k bootstrap CIs, three n-gram baselines, anchored 1–5 rubric over 60 prompts, automatic repetition metrics
- Deployment —
SolGeneratorinsrc/infer.py, Streamlit on Community Cloud, weights on Hugging Face Hub
Try it live — give it the opening of a children's story and it finishes it. Source on GitHub — clone it, run the tests, read the data card.
Skills this demonstrates
| Skill area | Evidence in the repo |
|---|---|
| ML engineering | Hand-written transformer, training loop, LR schedule, checkpoint/resume, M4 smoke gates |
| Data science | Dedup funnel, cross-split leakage fix, EDA, data card, bootstrap CIs, ablations with seed-variance yardstick |
| Evaluation design | Perplexity vs 3 baselines, anchored rubric (60 prompts), automatic repetition metrics |
| Systems / MLOps | YAML configs, spec drift CI test, 154 tests, export pipeline, deployment under 1 GB RAM |
| Honest communication | Limitations doc per milestone, null ablation result published, wrong pre-training target acknowledged |
Architecture and training
| Component | Choice | Why |
|---|---|---|
| Layers × heads | 8 × 8 (head_dim 74) | Fits 8 GB with headroom — peak 2,344 MiB measured |
| Embedding dim | 592 (not 512) | 512 measured only 41.8M params; 592 lands within 1.73% of the ~52M target |
| Context | 512 tokens | 99.74% of training documents fit whole — measured in EDA |
| Vocab | 32k BPE | Trained on train split only, byte-level |
| Precision | BF16 | Ada Lovelace native; no GradScaler needed |
| Throughput | ~20,067 tok/s | 40k iterations → 1.31B tokens over ~16.9 h awake training |
Attention uses scaled_dot_product_attention with causal masking (tested: perturbing future tokens cannot change logits at position t). Post-training, a KV cache was added with byte-identity proof at fp32 — 3.5× on CPU locally, 1.8× on the deployed Streamlit host.
Results
Validation perplexity
3.719
95% CI [3.693, 3.745], 15,141 documents, 10,000 bootstrap resamples
Peak VRAM
2,344 MiB
Target was < 7,400 MiB on an 8 GB RTX 4070 Laptop GPU, BF16, no gradient checkpointing
Vs baselines (same val set, same tokenizer)
Original target was 15–25 ppl — a pre-measurement guess. The interesting comparison is against baselines on identical data, not against that guess.
Anchored rubric
n = 60 prompts
Grammar scores 4.00/5 but coherence only 3.15/5 — clean sentences, but entity drift after ~150 tokens. Perplexity alone would not surface that gap.
Evaluated document-by-document over the full validation set, with bootstrap confidence intervals, against three baselines — because "perplexity 3.7" means nothing without knowing what a trigram model scores on the same data.
The pre-training target of 15–25 ppl was wrong, not the model exceptional. The interesting result is the rubric split: grammar strong, coherence weak. The model writes clean sentences and loses track of who is in the story.
What makes this not fine-tuning
| Sol | Typical fine-tuning project |
|---|---|
| Hand-written architecture with causality tests | Load pretrained weights |
| In-domain tokenizer trained on train split only | Reuse existing tokenizer |
| Dedup funnel + leakage detection + data card | Download dataset, minimal cleaning |
| Perplexity + bootstrap CI vs 3 baselines | "Loss went down" |
| Seed-variance yardstick for ablations | Single-run ablations |
Spec drift guard (export_spec.py + CI test) | Numbers in README drift |
| Deployment under 1 GB RAM constraint | gradio launch |
Data science focus
- Pipeline: 2,119,719 raw documents → 1,748,358 kept train docs / 357.9M tokens; documented dedup funnel tracing every number to its artifact.
- Leakage finding: 28.67% of raw validation docs were exact duplicates of training documents — not internal val redundancy. 6,304 dropped before any reported metric.
- Experimentation: Learning-rate sweep and data-scale ablation, both read against measured seed variance (±0.0045 ppl) rather than eyeballed.
- Honest null: Data scale (100M vs full 357.9M tokens) moved perplexity only ~18× the seed-noise floor — TinyStories' 14.58% within-train exact-dup rate may explain why a smaller slice still looks like the full corpus.
Ablations
Three identically-configured seeds: 4.490 ± 0.0045 perplexity — the yardstick for everything else.
- Learning rate (1e-4 / 3e-4 / 1e-3 → 5.380 / 4.489 / 4.162) — gaps ~271× the seed-noise floor. Large, real effect.
- Data scale (100M vs full corpus → 4.572 vs 4.489) — only ~18× the floor. The variable that sounds like it should matter most mattered least.
Deployment
Hugging Face moved Gradio Spaces behind PRO mid-project (402 Payment Required). The Gradio app was already working; porting to Streamlit Community Cloud took about an hour with zero inference code changes because SolGenerator lives in one module.
Engineering for the free tier: pin Python 3.12 and CPU-only torch, @st.cache_resource under a 1 GB RAM ceiling, 103.1 MiB weight bundle, KV cache bringing deployed generation from 16 → 29 tok/s.
Limitations
Coherence is the binding constraint — named entities drift across roughly 150 tokens. Children's stories only; no instruction tuning; will continue a question as prose rather than answer it (on-topic 5.00/5 in-domain vs 1.47/5 out-of-domain). Mojibake in a few percent of generations traces to 6.20% of upstream TinyStories documents — inherited, documented, not quietly patched.
All tracked in the repository's rolling limitations document, written per milestone rather than assembled at the end.
Highlights
- Hand-written decoder-only transformer (~52.9M params) — architecture, data pipeline, training, eval, ablations, and deployment; no pretrained weights
- Validation perplexity 3.719 (95% CI [3.693, 3.745]) vs trigram 23.4 on the same 15,141-document val set
- Cross-split leakage caught: 6,304 validation documents were exact duplicates of training docs — dropped before eval
- Grammar 4.00/5 but coherence 3.15/5 on an anchored rubric — perplexity alone would miss the real limit
- 154 tests + generated spec with CI drift guard; live demo on Streamlit free tier (103 MiB weights on Hugging Face)
Challenges
Outcomes
- Demonstrates end-to-end ML ownership: data → train → eval → ablation → deploy — not API integration
- M0–M9 roadmap complete on a single 8 GB laptop GPU (BF16, 1.31B tokens, 40k iterations)
- Public artifact recruiters can inspect: GitHub repo, Hugging Face weights, live Streamlit demo
- Interview-ready numbers with CI drift guard — every metric traces to a test or artifact
Demo
Open live appValidation perplexity
3.719
95% CI [3.693, 3.745] over 15,141 documents
vs trigram baseline
23.4
Unigram 379, uniform 32,000
Rubric: grammar
4.00 / 5
Hand-scored over 60 prompts
Rubric: coherence
3.15 / 5
The binding limitation — entity drift over ~150 tokens, not grammar
The original spec targeted perplexity 15–25. Measured 3.719 — which means the target was wrong about how hard TinyStories is, not that the model over-performed. Deployed inference runs at 29 tokens/second on a free shared vCPU.