The model learns general language and knowledge by predicting the next token across enormous amounts of text. This is the most expensive stage.
12
Beginner
Pt
Pretraining
Data & Training
Pretraining
The first, massive phase of training.