Notes

Knowledge notes

Smaller explanations, fragments, and references that can grow into full articles later.

Allconcurrencydeep-learningdotnetef-corelayernormtransformers
notesAug 6, 20261 min read

Pre-Layer Normalization in Transformers

A short note on why moving LayerNorm before attention and feed-forward blocks can improve training stability.

transformerslayernormdeep-learning