Residual connections and layer normalization

Residual (skip) connections and layer normalization are what make deep transformers trainable. Without them, gradients vanish or explode before reaching early layers. Pre norm (LayerNorm before attention/MLP) is now standard — it stabilizes training at the cost of some representational capacity.