Skip to Content
LLM Fundamentals II – Inside the ModelAn opened decoder-only model: input tokens, embeddings, repeated pre-norm Transformer blocks with causal attention and MLP residual paths, then final normalization and vocabulary projection for next-token probabilities.APPLIED AI FOR DEVELOPERSLLM Fundamentals IIInside the ModelIn Part I, the model was a black box.Now we open it up.Look insideFollow tokens throughthe decoder stack.Understand attentionSee how context shapeseach next-token prediction.Build better intuitionConnect architecture toeveryday development.DECODER-ONLY LANGUAGE MODELINPUTThefutureofAIisTokenembeddingsTransformer block × NNorm → causal attention+ residualNorm → MLP+ residualRoPE: position in Q and KCommon pre-norm designFinal norm+ vocabprojectionNEXT TOKENprobabilitiesIllustrativedistributionTODAY’S ROUTETokens → representations → attention → context → next-token prediction
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com1 / 16
0102030405060708The futureof AI is…The101future205of309AI412is518AttentionMLPChoose a tokenbrightThefutureofAIisbrightAppend the token, embed it, and predict againText(Input)TokenizationToken IDsTokenEmbeddings(Vectors)Model(Transformer)Next-TokenProbabilitiesSelect NextTokenAppend &RepeatOUTPUTGenerated Text(Output)The futureof AI isbright.Key Takeaway:LLMs don’t work on text directly – they operate on numbers (vectors) and predict the most likelynext token, one step at a time.Your promptExample token splitIllustrative IDsLearned vectorsSimplified stackIllustrativedistributionSampling orgreedy decodingContinue untila stop conditionStop conditionhas been met.× N blocksAPPLIED AI FOR DEVELOPERSRecap:From Text to PredictionIn Part I, we saw the end-to-end flow of how an LLM turns your text into the next token.Here’s the big picture again.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com2 / 16
APPLIED AI FOR DEVELOPERSWhat Is Inside theBlack Box?A decoder-only Transformer turns token IDs into contextual vectors, then next-token scores (logits).Shown here: a common sequential pre-norm design, as used in Llama-style models.TEXT → TOKENS → IDs“The future of AI is …”The101future205of309AI412is518Illustrative split and IDsINSIDE THE MODEL1 EmbeddinglookupOne vector per token2 Decoder block × NSame structure; separate learned weightsNormCausal self-attentionNormMLP / feed-forward++3 Final normContextualrepresentationsLast position →4 Vocabulary headLinear projection → logitsbright3.2changing1.6unclear−0.7Illustrative unnormalized scoresSoftmax → probabilitiesThen choose the next tokenRotary Position Embedding (RoPE) rotates query and key vectors according to token position.This makes attention sensitive to relative position. The causal mask blocks future tokens.KEY IDEAAttention mixes context; MLPs transform features; the vocabulary head scores next tokens.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com3 / 16
APPLIED AI FOR DEVELOPERSWhyContext MattersThe same token can have different contextual representations and next-token predictions.In a causal decoder, each position uses itself and earlier tokens — never future tokens.A FINANCIAL CONTEXTEarlier contextTodepositcash,IvisitedthebankCausal attention + MLPs across layersContextual representation of “bank”Illustrative vectorPossible next tokens:branchtoand“bank” refers to a financial institution.Examples, not measured rankingsB RIVERSIDE CONTEXTEarlier contextAlongtheriver,wesatonthebankCausal attention + MLPs across layersContextual representation of “bank”Illustrative vectorPossible next tokens:andofnear“bank” refers to the side of a river.Examples, not measured rankingsKEY IDEAA token’s identity stays the same; its representation changes with the preceding context.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com4 / 16
APPLIED AI FOR DEVELOPERSThe Transformer –The Core ArchitectureMany modern LLMs use a decoder-only Transformer: repeated blocks refine contextual representations.A common pre-norm design is shown below. Exact block, position and attention choices vary by model.WHY THIS ARCHITECTURE?Context mixingCausal attention combinesinformation from allowed tokens.Depth + residual pathsLayers refine features; skipconnections support training.Parallel computationTraining / prefill: positionsin parallel. Decoding: stepwise.INPUTDECODER STACKOUTPUTThe future of AI isTokenized promptEmbedding lookupONE DECODER BLOCK · REPEAT N TIMESNorm → causal self-attention+ residualNorm → MLP+ residualEach block has its own learned weightsRoPE supplies position within attentionFinal norm → vocabulary projectionUse the last position to score the next tokenLogitsbright 3.2 new 1.1Illustrative raw scoresSoftmaxToken probabilitiesSelect: “bright”Sampling or greedy selection;then append and repeat.Embedding vectors → stackAttention mixes positions. MLPs transform features. Norm and residual paths stabilize the stack.KEY IDEARepeated decoder blocks build context; the output head turns it into next-token scores.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com5 / 16
APPLIED AI FOR DEVELOPERSThe TransformerBlockOne decoder block contains two sublayers: attention mixes context; the MLP transforms token features.Sequential pre-norm design: normalize before each sublayer, then add its update to the residual stream.ONE DECODER BLOCK · TWO SUBLAYERSxNorm1 CAUSAL SELF-ATTENTIONCombine information from allowed token positions+hResidual / skip pathhNorm2 POSITION-WISE MLPTransform the features of each token separately+yResidual / skip pathPosition information and causal masking guide attention.Same MLP weights at each position; no mixing across positions here.WHAT EACH PART DOESAttentionEach position uses itself and earlierpositions to combine contextual features.Norm + residualNormalize the sublayer input; add itsoutput to the unchanged residual path.MLP / feed-forwardTransforms features independentlyat each position.Norm = normalization. MLP = multi-layer perceptron. Exact normalization and activation choices vary by model.PRE-NORMh = x + Attention(Norm(x))y = h + MLP(Norm(h))
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com6 / 16
APPLIED AI FOR DEVELOPERSAttention –The Core IdeaFor each token position, attention gathers a weighted mix of information from accessible positions.The weights depend on the input, the layer and the attention head. They are computed anew for each query.EXAMPLE · PROCESSING “bank”Todepositcash,IvisitedthebankyesterdayEarlier tokens + the current token: allowedFuture: maskedAttention weights for the query at “bank”0.040.340.280.030.100.060.150.00High weight → a larger contribution from that position’s value vector.Illustrative weights for one head; not measured model outputs. Weights sum to 1.HOW THE MIX IS COMPUTED1 CompareThe current query is compared withthe keys of accessible positions.2 NormalizeMask future positions, then use softmaxto turn the scores into weights.3 MixMultiply each value vector by its weightand sum them into one output vector.Queries, keys and values are learned projections of token representations. The next slides unpack them.KEY IDEAAttention mixes value vectors — it does not select the next token or prove which words explain an answer.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com7 / 16
APPLIED AI FOR DEVELOPERSSelf-Attentionin PracticeSlide 7 showed one query at “bank”. Now apply the same operation at every token position.Each row has its own query, allowed keys and attention weights. All positions share the learned projections.WHO CAN ATTEND TO WHOM?Query ↓ / Key →Todepositcash,IvisitedthebankyesterdayTo×××××××deposit××××××cash,×××××I××××visited×××the××bank.04.34.28.03.10.06.150yesterdayBlue: allowed · Gray: masked · Purple row: the illustrative weights from slide 71 One query per position“bank” can use itself and earlier tokens.“yesterday” can also use “bank”.“bank” cannot look ahead to “yesterday”.2 One contextual output per positionEach row mixes the accessible value vectors.The purple row sums to 1; the other rowshave their own weights, not shown here.3 Parallel computation, causal accessTraining and prompt prefill compute rowsin parallel. The mask still prevents future access.“Self” means queries, keys and values come from the same sequence. Diagram: one head in one layer.KEY IDEAEvery position builds its own contextual representation — without seeing future tokens.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com8 / 16
APPLIED AI FOR DEVELOPERSQueries, Keysand ValuesThree learned linear projections produce queries, keys and values from the same input representations.Continue the “bank” example: its query is compared with accessible keys, then the resulting weights mix values.ONE INPUT · THREE LEARNED PROJECTIONSX: one normalized token vector per rowQUERYQ = X · WQKEYK = X · WKVALUEV = X · WVQuery row at “bank”K at allowed positionsV at allowed positionsQ · K → scaled scores → mask → softmaxAttention output = weighted sum of V vectorsQuery: what matches this position?For “bank”, Q determines how it scoresthe available keys in this attention head.Key: how does a position match?Each position supplies a K vector.With RoPE, Q and K rotate before scoring.Value: what information is mixed?Each position also supplies a V vector.The weights scale these vectors beforethey are summed into one output.One head shown. X contains one row per token; W matrices are learned projections shared across positions.KEY IDEAQueries and keys determine the weights. Values supply the information being combined.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com9 / 16
APPLIED AI FOR DEVELOPERSMulti-HeadAttentionSeveral attention heads process the same sequence in parallel, using different learned projections.Their outputs are joined, projected and passed to the residual addition in the decoder block.ONE SEQUENCE · SEVERAL PARALLEL HEADSX: normalized token representationsHEAD 1Own Q, K and V projectionsCausal attention → z1Same past-token limitHEAD 2Own Q, K and V projectionsCausal attention → z2Same past-token limitHEAD 3Own Q, K and V projectionsCausal attention → z3Same past-token limitConcatenate [z₁ | z₂ | z₃]WₒYWHAT THE DIAGRAM MEANS1 Each head has learned Q/K/V projections.It can assign different attention weights.2 Outputs are concatenated, then mixedby a learned output projection Wₒ.3 Heads have no fixed semantic labels.Patterns and useful roles emerge in training.MODEL VARIANTSMHA: separate K/V per query head.GQA / MQA: K/V shared by groups / all heads.Three heads are shown only for clarity. Head counts and K/V sharing depend on the model configuration.KEY IDEAMultiple learned views of the same context are combined into one attention output.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com10 / 16
APPLIED AI FOR DEVELOPERSPositionMattersThe same words in a different order can describe a different event.Many decoder models use rotary position embeddings (RoPE) to make attention position-sensitive.SAME WORDS · DIFFERENT ROLES012345POSThedogchasedthecat.Chaser: dogChased: catThecatchasedthedog.Chaser: catChased: dogWord boxes are illustrative; a tokenizer may split words into smaller pieces.RoPE · ROTARY POSITION EMBEDDINGFor one attention head: rotate Q and K by position.QUERY at irotate by iKEY at jrotate by jcomparerotated Q · rotated KRelative distance influences the score.The causal mask separately blocks future positions.RoPE is one position method (used by Llama). Other models use learned positions or attention biases.KEY IDEAPosition helps attention distinguish both token content and where it occurs.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com11 / 16
APPLIED AI FOR DEVELOPERSLayer byLayerPart I gave us token embeddings: learned starting vectors for token IDs.Decoder blocks repeatedly update the sequence, making each position sensitive to its available context.FROM TOKEN EMBEDDING TO CONTEXTUAL REPRESENTATIONExample prompt · words shown for clarityIdepositedcashatthebank.bank = focusPART IEmbedding lookupsame token ID → same startBLOCK 1Attend + transformearlier tokens contributeBLOCK 2Attend + transformfeatures are updatedBLOCK NAttend + transformready for predictionOnly “bank” is highlighted here. Every position passes through every block; each block has its own weights.WHAT EACH BLOCK DOES1 Causal self-attentionMixes information from earlier and current positions.2 MLP / feed-forwardTransforms features at each position.3 Residual paths + pre-normCarry the stream forward around both sublayers.Then: final norm → vocabulary scores (logits).Layer roles are not fixed: there is no universal “grammar layer” or “facts layer”.KEY IDEAEmbeddings start the sequence; stacked blocks build context-sensitive states for prediction.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com12 / 16
APPLIED AI FOR DEVELOPERSContextWindowThe model uses the tokens available for this request, including relevant earlier messages when supplied.The application decides what to send; chat roles and content are encoded in a model-specific token sequence.WHAT MAY BE IN THE CURRENT CONTEXT?System / developer instructionsif usedEarlier user + assistant turnsif retainedTool results / retrieved documentsif suppliedCurrent user messagenew inputAssistant tokens generated so farduring outputFINITETOKENBUDGETOlder items can be omitted or summarized when the window fills.The exact limit depends on the model and serving configuration.CONTEXT WINDOW ≠ LONG-TERM MEMORYA window is working input for this request.Prompting does not itself change model weights.Saved facts return only if the app includes them.FOR DEVELOPERS1Budget for outputInput and generated tokens share the limit.2Select useful contextRetrieve relevant text; summarize old turns.3Check what survivesTest truncation and verify critical facts.A larger window holds more tokens but does not guarantee that every detail will be used reliably.KEY IDEAYour app must select, fit and refresh the information the model needs.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com13 / 16
APPLIED AI FOR DEVELOPERSTrainingvs InferenceThe same decoder model is used to learn from known sequences or to generate from a prompt.Training compares predictions with known next tokens; inference extends a prompt with new tokens.TRAINING · LEARN WEIGHTSKnown sequence: “The future of AI is bright”INPUT POSITIONSThefuturefutureofofAIAIisisbrightKNOWN NEXT-TOKEN TARGETSCausal forward pass → scores at all five positionsCompare with targets → loss → backprop → weight updateEach position sees its left prefix; many losses run in parallel.INFERENCE · USE FIXED WEIGHTSPrompt: “The future of AI is”ThefutureofAIisPrefill: process the prompt positionsNext-token scores → choose one tokenAppend “bright” → decode the next stepOutput tokens arrive sequentially; weights stay fixed.A KV cache can reuse earlier keys and values.Word boxes simplify tokenization; “bright” is an example, not a guaranteed prediction.KEY IDEATraining updates weights from errors; inference generates with fixed weights, one new token at a time.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com14 / 16
APPLIED AI FOR DEVELOPERSWhy LLMsCan Be WrongA fluent answer can contain unsupported or incorrect claims.Next-token prediction produces text; factual accuracy needs separate evidence and checks.WHERE ERRORS CAN ENTERQuestion + contextModel scores tokensFluent answerFLUENT ≠ VERIFIEDA claim needs support beyond how likely its words sound.1Learned patternsTraining text can include mistakes or misconceptions.2Missing evidenceThe needed facts may be absent, old or poorly used.3Reasoning / generationA calculation, inference or multi-step answer can fail.Different tasks and models fail in different ways; test your own use case.BUILD VERIFIABLE ANSWERS1Ground in evidenceRetrieve relevant, trusted source material.2Check citationsMatch each important claim to the actual source.3Use toolsValidate calculations, IDs and structured facts.4Evaluate + abstainTest hard cases; allow “I don’t know”.A next-token probability is not the probability that a factual statement is true.KEY IDEATreat generated claims as proposals: verify them against trusted evidence and tools.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com15 / 16
APPLIED AI FOR DEVELOPERSKey Takeaways& What’s NextWe have followed tokens from embeddings through decoder blocks to next-token scores.Use this mental model—and its limits—when designing AI applications.THE MODEL IN FOUR IDEAS01 TOKENS → VECTORSPart I: token IDs map to learned embeddings.These are starting vectors for the prompt.02 LAYERS BUILD CONTEXTCausal attention and MLPs update vectors.Norms and residual paths support the stack.03 PREDICT & REPEATFinal states produce next-token logits.Choose a token, append it, then continue.04 KNOW THE LIMITSContext is finite; fluent text can be wrong.Evidence and verification remain essential.Position methods, head sharing and block details vary by model; inspect its configuration.WHAT’S NEXT FOR DEVELOPERS1Inspect a modelLayers, heads, position method, context limit.2Trace one promptToken IDs → logits → selected output tokens.3Build for reliabilityAdd relevant evidence, tools and validation.Then evaluate on real tasks and failure cases.Decoder-only models share core ideas, but implementations are not identical.KEY IDEAUnderstand the mechanism; then design the context and checks around it.
LLM FUNDAMENTALS II · 2026Dr. Jürgen Nützel, www.deenovum.com16 / 16

Arrow keys: navigate · Home / End: first / last slide · Download PDF: downloads all 16 slides directly, without a print dialog. The PDF is a saved edition; regenerate it after changing slide content.