← Back to KHAO

Llama ·

Tuning a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

Eval loss and CORE for the 1024 and 2048 context runs.

Somewhere between “nanoGPT toy” and “you need a research lab” there’s a large, under-described region where one person with a few thousand dollars can train a meaningful model.

Key facts

Summary

The reporter wanted to see language and understanding emerge from random weights for themselves, and to learn the parts you can only learn by starting from scratch. The result is a 3.8B-parameter model scoring 0.384 on CORE, trained on 65B tokens in 43 hours for $998. What follows is what worked, what didn’t, and what the reporter still don’t know. Their model is larger than nanochat d32 and took similar wall-clock time. But for roughly the same money as nanochat’s $1,000 configuration, this lands meaningfully ahead of it.

Read full article at hugovergnes.github.io →

#Llama #Nvidia B200