Why the textbook XOR swizzle still bank-conflicts a 16×32 transpose — and how the optimal one fixes it with a single bit shift.