What are LLMs mainly trained on?
AnswerLarge amounts of text
They learn patterns of language from huge text collections.
AI · Facts & stats
7 fact-checked facts about large language models and how they work, each with the reason behind it. 5 more are in today's round and join this page after it closes.
AnswerLarge amounts of text
They learn patterns of language from huge text collections.
AnswerThe input you give the model
A prompt is the text or instruction the model responds to.
AnswerChunks of text the model processes
Text is split into tokens, which may be words or word pieces.
AnswerGenerative Pre-trained Transformer
GPT models are pre-trained on text and generate new text.
AnswerReinforcement learning from human feedback
RLHF uses human ratings to steer model behavior.
AnswerMakes output more random
Higher temperature spreads probability across more word choices.
AnswerWord embeddings (vectors)
Word2vec, from 2013, maps words to vectors that capture meaning.