All work
Done2026

Sol

Applied ML skills portfolio piece — train, measure, ablate, and ship a language model from scratch

Role
ML engineer
Stack
PyTorchTinyStoriesBPE TokenizationStreamlitHugging Face HubCUDAYAML Configs
Training and validation cross-entropy loss for Sol's 52.9M-parameter baseline over 40k iterations, both settling near 1.3 nats
Open live appSource

Sol's demo sits behind a login — the app generates text on every request, and an open URL would let anyone drain its budget. Get in touch and I'll send you credentials. Request credentials

Why I built this

Sol is not a product trying to solve a user problem — it is a skills demonstration for applied ML / ML engineering roles. Recruiters and hiring managers evaluating me need evidence that I understand training dynamics, tokenization, evaluation design, and hardware trade-offs — not just how to call an API.

Fine-tuning a pretrained model is faster but weaker signal for fundamentals. Sol is the opposite: a small transformer trained properly and measured honestly, with every number guarded by tests and artifacts, so you can see how I work rather than what I shipped.

What it is

A 52,901,712-parameter decoder-only transformer, written by hand in src/model.py (~300 lines, nanoGPT-informed but not forked) and trained from scratch on TinyStories on a single RTX 4070 Laptop GPU (8 GB, BF16). No pretrained weights. The repository is the résumé — it contains the full stack:

  • Data — download, dedup funnel, cross-split leakage detection, in-domain 32k byte-level BPE, binary shards, EDA notebook, data card
  • Training — cosine LR schedule, grad clip, checkpoint/resume with RNG state, M4 smoke gates (overfit-one-batch, VRAM benchmark)
  • Evaluation — full-val perplexity with 10k bootstrap CIs, three n-gram baselines, anchored 1–5 rubric over 60 prompts, automatic repetition metrics
  • Deployment — SolGenerator in src/infer.py, Streamlit on Community Cloud, weights on Hugging Face Hub

Try it live — give it the opening of a children's story and it finishes it. Source on GitHub — clone it, run the tests, read the data card.

Skills this demonstrates

Skill areaEvidence in the repo
ML engineeringHand-written transformer, training loop, LR schedule, checkpoint/resume, M4 smoke gates
Data scienceDedup funnel, cross-split leakage fix, EDA, data card, bootstrap CIs, ablations with seed-variance yardstick
Evaluation designPerplexity vs 3 baselines, anchored rubric (60 prompts), automatic repetition metrics
Systems / MLOpsYAML configs, spec drift CI test, 154 tests, export pipeline, deployment under 1 GB RAM
Honest communicationLimitations doc per milestone, null ablation result published, wrong pre-training target acknowledged

Architecture and training

ComponentChoiceWhy
Layers × heads8 × 8 (head_dim 74)Fits 8 GB with headroom — peak 2,344 MiB measured
Embedding dim592 (not 512)512 measured only 41.8M params; 592 lands within 1.73% of the ~52M target
Context512 tokens99.74% of training documents fit whole — measured in EDA
Vocab32k BPETrained on train split only, byte-level
PrecisionBF16Ada Lovelace native; no GradScaler needed
Throughput~20,067 tok/s40k iterations → 1.31B tokens over ~16.9 h awake training

Attention uses scaled_dot_product_attention with causal masking (tested: perturbing future tokens cannot change logits at position t). Post-training, a KV cache was added with byte-identity proof at fp32 — 3.5× on CPU locally, 1.8× on the deployed Streamlit host.

Results

Validation perplexity

3.719

95% CI [3.693, 3.745], 15,141 documents, 10,000 bootstrap resamples

Peak VRAM

2,344 MiB

Target was < 7,400 MiB on an 8 GB RTX 4070 Laptop GPU, BF16, no gradient checkpointing

Vs baselines (same val set, same tokenizer)

Sol (transformer)3.719
Trigram23.4
Unigram379
Uniform32,000

Original target was 15–25 ppl — a pre-measurement guess. The interesting comparison is against baselines on identical data, not against that guess.

Anchored rubric

n = 60 prompts

Grammar4.00 / 5
Coherence3.15 / 5
On-topic (in-domain)5.00 / 5
On-topic (out-of-domain)1.47 / 5
Repetition3.98 / 5

Grammar scores 4.00/5 but coherence only 3.15/5 — clean sentences, but entity drift after ~150 tokens. Perplexity alone would not surface that gap.

Evaluated document-by-document over the full validation set, with bootstrap confidence intervals, against three baselines — because "perplexity 3.7" means nothing without knowing what a trigram model scores on the same data.

The pre-training target of 15–25 ppl was wrong, not the model exceptional. The interesting result is the rubric split: grammar strong, coherence weak. The model writes clean sentences and loses track of who is in the story.

What makes this not fine-tuning

SolTypical fine-tuning project
Hand-written architecture with causality testsLoad pretrained weights
In-domain tokenizer trained on train split onlyReuse existing tokenizer
Dedup funnel + leakage detection + data cardDownload dataset, minimal cleaning
Perplexity + bootstrap CI vs 3 baselines"Loss went down"
Seed-variance yardstick for ablationsSingle-run ablations
Spec drift guard (export_spec.py + CI test)Numbers in README drift
Deployment under 1 GB RAM constraintgradio launch

Data science focus

  • Pipeline: 2,119,719 raw documents → 1,748,358 kept train docs / 357.9M tokens; documented dedup funnel tracing every number to its artifact.
  • Leakage finding: 28.67% of raw validation docs were exact duplicates of training documents — not internal val redundancy. 6,304 dropped before any reported metric.
  • Experimentation: Learning-rate sweep and data-scale ablation, both read against measured seed variance (±0.0045 ppl) rather than eyeballed.
  • Honest null: Data scale (100M vs full 357.9M tokens) moved perplexity only ~18× the seed-noise floor — TinyStories' 14.58% within-train exact-dup rate may explain why a smaller slice still looks like the full corpus.

Ablations

Three identically-configured seeds: 4.490 ± 0.0045 perplexity — the yardstick for everything else.

  • Learning rate (1e-4 / 3e-4 / 1e-3 → 5.380 / 4.489 / 4.162) — gaps ~271× the seed-noise floor. Large, real effect.
  • Data scale (100M vs full corpus → 4.572 vs 4.489) — only ~18× the floor. The variable that sounds like it should matter most mattered least.

Deployment

Hugging Face moved Gradio Spaces behind PRO mid-project (402 Payment Required). The Gradio app was already working; porting to Streamlit Community Cloud took about an hour with zero inference code changes because SolGenerator lives in one module.

Engineering for the free tier: pin Python 3.12 and CPU-only torch, @st.cache_resource under a 1 GB RAM ceiling, 103.1 MiB weight bundle, KV cache bringing deployed generation from 16 → 29 tok/s.

Limitations

Coherence is the binding constraint — named entities drift across roughly 150 tokens. Children's stories only; no instruction tuning; will continue a question as prose rather than answer it (on-topic 5.00/5 in-domain vs 1.47/5 out-of-domain). Mojibake in a few percent of generations traces to 6.20% of upstream TinyStories documents — inherited, documented, not quietly patched.

All tracked in the repository's rolling limitations document, written per milestone rather than assembled at the end.

Highlights

  • Hand-written decoder-only transformer (~52.9M params) — architecture, data pipeline, training, eval, ablations, and deployment; no pretrained weights
  • Validation perplexity 3.719 (95% CI [3.693, 3.745]) vs trigram 23.4 on the same 15,141-document val set
  • Cross-split leakage caught: 6,304 validation documents were exact duplicates of training docs — dropped before eval
  • Grammar 4.00/5 but coherence 3.15/5 on an anchored rubric — perplexity alone would miss the real limit
  • 154 tests + generated spec with CI drift guard; live demo on Streamlit free tier (103 MiB weights on Hugging Face)

Challenges

Outcomes

  • Demonstrates end-to-end ML ownership: data → train → eval → ablation → deploy — not API integration
  • M0–M9 roadmap complete on a single 8 GB laptop GPU (BF16, 1.31B tokens, 40k iterations)
  • Public artifact recruiters can inspect: GitHub repo, Hugging Face weights, live Streamlit demo
  • Interview-ready numbers with CI drift guard — every metric traces to a test or artifact

Demo

Open live app

Validation perplexity

3.719

95% CI [3.693, 3.745] over 15,141 documents

vs trigram baseline

23.4

Unigram 379, uniform 32,000

Rubric: grammar

4.00 / 5

Hand-scored over 60 prompts

Rubric: coherence

3.15 / 5

The binding limitation — entity drift over ~150 tokens, not grammar

The original spec targeted perplexity 15–25. Measured 3.719 — which means the target was wrong about how hard TinyStories is, not that the model over-performed. Deployed inference runs at 29 tokens/second on a free shared vCPU.

Let's work together

manuel@manuelvargas.dev

Open to full-time roles and freelance contracts. I reply within 48 hours.