Generative AI Basics - Part 1

published Sep 28, 2026 3:08pm

This series of posts is intended to introduce the reader to the mechanics of how generative AI tools (e.g. ChatGPT, Gemini, Claude) actually work under the hood. The goal is to provide enough detail to allow the reader to build a reasonably accurate conceptual understanding so that when they are making decisions about how to use these tools, those decisions can be informed by how they actually work. We will gloss over lots of details (and entire fields of study) that are very interesting but not critical to building that conceptual understanding. But we will also take the time to focus specific attention on small details when they are critical to making the big picture hang together coherently. So if we seem to be getting into the weeds, do your best to hang with us. We assume the reader is a geoscience educator and is comfortable with the concepts (especially mathematical) that a typical geoscientist would be familiar with. But we hope this will be approachable more broadly; at the very least, to STEM educators in other fields.

Models (in general and more specific)

We will focus our discussion on large language models (LLMs). These are the core piece of the generative AI tools mentioned above, and contain the main bits of magic whose workings we want to expose. From one perspective, these are very much like many other numeric models we might build to do useful work. We might use a weather model to predict the temperature tomorrow in a particular location, or a hydrological model to predict the flow rate in a particular stream on a particular day.  In both these (and many other) models we have some sort of numerical input: the date and location of our desired predicted temperature and data about the state of the atmosphere currently, the location in the stream and date for our flow prediction, as well as other hydrological input that may influence the prediction like recent rainfall amounts. These inputs are then fed through some sequence of mathematical operations which combine our inputs with other numbers. The 'other' numbers are numeric parameters, often derived from earlier scientific work, and may represent physical properties (the mass of a gram of water), or reflect the dynamics of the system (the entrainment rate of cloud formation or Manning's Roughness Coefficient for stream flow).  At the end of all the math, we get numeric output: the model's predictions for temperature or streamflow.  We can run the model repeatedly with different inputs. Each time the model repeats the same sequence of calculations with the same internal parameters, but with changed inputs. The result is new output numbers (temperature, streamflow) that correspond to our new inputs. A model is good if the outputs for a given input are usefully close to what happens in the reality they are trying to model.

LLM's follow a similar design. There are user-chosen inputs and a set of fixed numerical parameters 'inside' the model. The model itself is just a sequence of mathematical operations (mostly additions and multiplications with a few square roots and exponentiations) that intermixes the inputs and parameters and eventually spits out a set of output numbers. You can put in different inputs and get different outputs. Though, because the mathematical steps and parameters are fixed, the same inputs always give the same outputs for a given model. (If you find this statement surprising, don't worry, we'll come back to it in part 2).  The input is always a string of text that may contain words and spaces and punctuation. It could be a short, even incomplete, string of text The elephant has a long or an extended run of text (e.g. all the words on page 270 of whatever textbook you have at hand).  The output is a set of numbers which predict the very next word that should occur after your input string.  In our elephant example, those numbers might include the prediction that trunk has a 93% probability of being the next word. In our textbook example, it might predict the word igneous has a 53% chance of being the next word.

Tokens - the Input

Since LLM's are a set of mathematical operations we first need to transform your string of text into a set of numbers.  We do this by pre-assigning specific numbers to all possible short words and word fragments. Then we can just look up the sequence of numbers that match the sequence of words you typed in.  So, for instance, the (incomplete) phrase The elephant has a long is represented (in ChatGPT for example) by a series of 5 numbers: 855, 28212, 1109, 304, 1583.  The first word ('The' with a capital T) is represented by 855.  28212 represents a single space followed by the word 'elephant'. The other numbers similarly represent the rest of the words in the phrase and their preceding spaces.  Imagine a big spreadsheet with a column for words (with spaces around them or not) and a column for their corresponding number.  These numbers are called 'tokens' and ChatGPT has a fixed set of 200,000 tokens which cover all the possible words and punctuation it might encounter. Note that there is some 'cheating' in this tokenization as very long words may be represented by several tokens in a row. For example the word unbelievable is represented by two tokens in a row: 68 (un) and 65851 (believable).  You can imagine that by wisely devising your token scheme to include parts of words like 'un' and 'ing' you can use a small library of tokens and still be able to represent any possible word simply by stringing several tokens together. And that's exactly what happens. ChatGPT's library of 200,000 tokens is sufficient to represent any set of words in English (and in fact it covers most written languages).  The details of how this tokenization happens aren't critical; different LLM's use different specific token mappings, but the end result is the same. The run of text is turned into a sequence of numbers and that's the input to our model.  You'll encounter the word 'token' when dealing with LLM's (e.g. pricing is sometimes 'per token') so it's good to know that it's a pretty simple concept. It's just an arbitrary mapping of words (plus word parts and punctation) to a set of numbers.

Next Token Probabilities - the Output

The output of the model is a list of probabilities. It is literally just a list of numbers, each between 0 and 1 reflecting the probability of the very next token.  In our ChatGPT case there are 200,000 different probabilities that get spit out: one for every token that it knows about.  So, if we input our elephant phrase from above, and we looked through this output we might see an entry like '855, 0.0001', indicating that token 855 has a 0.01% chance of being the next word.  That makes sense because we saw above that token 855 was the word 'The' with no surrounding spaces and the phrase The elephant has a longThe seems pretty unlikely.  But if we dug further into those 200,000 probabilities we might find an entry like '32120, 0.74' indicating token 32120 (which is the token for trunk with a space in front of it) has a 74% chance of being the next token.  That would make our predicted phrase The elephant has a long trunk which seems more plausible.

So that's the sum total of what an LLM can do. You put in a string of text, it maps it to tokens and spits out probabilities for the next token that ought to appear. Of course we've left out the interesting bit; how does it make this prediction?  We'll get there next, but it's worth holding in your head that there's nothing beyond this 'next token prediction' in the LLM itself. Of course, your own experience with AI tools is probably that they do more than give you probabilities for a predicted single next word.  A mystery we will tackle further on.

The Math Inside the Model

Of course, in any good model the interesting bits are in the math inside.  We use our understanding of the natural world to come up with some set of interconnected equations that translate our input numbers into meaningful output values. So you might imagine that in order to construct an LLM, folks have invented or derived some complicated numerical models that reflect the logic of language in some way that allows the input numbers to flow to appropriate outputs.  This has indeed been tried.  But language is complicated. And to make the tool really useful, the equations have to somehow capture not only things like grammar (e.g. predicting trunk is better than predicting The because a noun is better than an article in that position), but also some knowledge of the world (e.g. predicting trunk is better than predicting antenna because that's how elephants are typical constructed).  So, when people have tried to manually construct mathematical models that capture all of that nuance, they have not had great success. With LLM's we do something different.

Rather than sweating the details about exactly what equations and parameters should be in our model, we start with a generic set of equations and some (literally) random parameters. The parameters are numbers between -1 and 1 (like 0.52 or -.03). The equations are a series of matrix multiplications interspersed with some non-linear rescaling operations.  They combine our input numbers with the internal parameters.  In the ChatGPT case we organize our equations so that they always output 200,000 numbers between 0 and 1 which we intend to interpret as the probabilities of the next word.  Or more precisely, our next token.  We turn our words into tokens, feed them into the set of equations with our initial random parameters, and out the other end we get....random junk.  The output 'probabilities' bear no relation to reasonable predictions.

This is not at all surprising, since after all, it's just random parameters and generic equations.  But then comes the clever bit.  We 'train' the model.  We do that using example text where we know the entire appropriate string of words. We take a string of words like The elephant has a long trunk,  and we feed the leading bit (without the trunk) into our model. Then we look at the predicted probabilities that come out and ask the critical question: what set of small changes might we make to our current (randomly set) parameters to make the output slightly more correct. In this case, how can we change the parameters to make the predicted odds for trunk, token number 32120 in our set of 200,000, just a little bit higher.  Our parameters become metaphorical knobs that we can adjust to change the output.  Because we already know a good output (trunk) we can figure out how to adjust the parameters in the right direction.  The math for determining these adjustments turns out to be just a bunch of derivatives. We know the equations in the model, so we can work out exactly how any parameter shift might impact the output.

Can That Actually Work?

At this point you should be skeptical. It seems implausible that any amount of knob twiddling (parameter adjusting) is going to get us to a place where our model does anything really useful. After all, our equations aren't intentionally organized in any way that would be helpful. Perhaps you can imagine adjusting the parameters so that the model always spits out trunk no matter the input.  But that's not useful.  But what if we made the model big? What if we made really, really big matrices with lots and lots of parameters? So that we had lots of knobs we could turn. And what if we took a whole bunch of word sequences where we knew the good 'next' token and turned the knobs in whatever direction made all the token predications just a little bit better?  And what if we did that repeatedly with lots of example text, each time turning the parameters a little further in 'good' directions? And what if we automated the whole thing and let it run and run?  Would we eventually get to a system, a set of parameters that, when we put in a new string of text, the prediction numbers coming out the other end made sense?  That made sense even if we fed the model a string of text that wasn't part of the training?  Our intuitive sense is probably to say no.  Perhaps we might hedge with caveats like, "well, maybe if it had as many parameters as there are atoms in the universe",  but not realistically. Language and all the understanding that goes into being able to predict a next word is just too complicated. But you'd be wrong. This is the point where the magic sits.  If you start out with enough random parameters (a few billion is sufficient) and you 'train' it on enough text (a few trillion tokens where you know the 'good' next words) you get an LLM that can predict reasonable next words for pretty much any input.

Now what exactly is going on that allows those particular parameters combined via those equations to give useful predictions is still sort of a mystery. And there's a whole interesting field on AI 'interpretability' which digs into details to try and figure out what's going on.  But the key thing to understand is that the ability to make useful predictions doesn't arise from people cleverly 'modeling' human linguistic ability with equations in the way that the equations in a weather model might be intentionally arranged to reflect what we know about the relevant physics.  It's an emergent property that is just the result of feeding lots of text strings in and letting the math (we're talking derivatives here) decide how to gradually adjust the parameters so they get better and better at producing known good values.  Once you've done that enough, and you have good parameters in hand, you can then use that (now fixed) set of parameters with new inputs and get useful output.  Ones that predict likely next words that seem to make sense.

It's worth sitting for a second and contemplating how you might, as a human faced with a new string of words, go about deciding what a good next word would be. What knowledge and understanding would you bring to bear? Then think about responses you've seen from AI tools and contemplate what might be going on inside the LLM as it does its prediction. Could it just be repeating word patterns it has seen before?  The math doesn't work for this.  Models are too small, in terms of parameters, to somehow directly encode the literal text it was trained on.  The parameters and equations must somehow model grammar and knowledge about the world in a more compact way.  But that's a whole interesting field of study that we will set aside in order to get to the next practical question.  How does this all relate to the AI tools I actually use? Because while I've certainly typed strings of text into an AI tool, I've never had one spit 200,000 probabilities back at me.  Let's continue on to part 2.

 



Generative AI Basics - Part 1 -- Discussion  

Join the Discussion


Log in to reply