Time the forward, backward, and optimizer step for the model sizes described in Section 2.1.2. Use 5 warmup steps and compute the average and standard deviation of timings over 10 measurement steps. How long does a forward pass take? How about a backward pass? Do you see high variability across measurements, or is the standard deviation small?
One caveat of benchmarking is not performing the warm-up steps. Repeat your analysis without the warm-up steps. How does this affect your results? Why do you think this happens? Also try to run the script with 1 or 2 warm-up steps. Why might the result still be different?
What is the total time spent on your forward pass? Does it match what we had measured before with the Python standard library?
forward 中大致都是 48 ms(除了 step 2),和之前的同济保持一致
What CUDA kernel takes the most cumulative GPU time during the forward pass? How many times is this kernel invoked during a single forward pass of your model? Is it the same kernel that takes the most runtime when you do both forward and backward passes? (Hint: look at the “CUDA GPU Kernel Summary” under “Stats System View”, and filter using NVTX ranges to identify which parts of the model are responsible for which kernels.)
选择 forward 区间并添加 filter,ampere_sgemm_128x64_tn 内核占用时间最长,这是一个矩阵运算 kernel,ampere 是安培架构(因为 GPU 用的是 3090),sgemm 是 Single precision (FP32) General Matrix Multiplication,tn 表示矩阵转置。
如果把 forward 和 backward 全算进去,同样这个 kernel 占用时间最长。
Although the vast majority of FLOPs take place in matrix multiplications, you will notice that several other kernels still take a non-trivial amount of the overall runtime. What other kernels besides matrix multiplies do you see accounting for non-trivial CUDA runtime in the forward pass?
Profile running one complete training step with your implementation of AdamW (i.e., the forward pass, computing the loss and running a backward pass, and finally an optimizer step, as you’d do during training). How does the fraction of time spent on matrix multiplication change, compared to doing inference (forward pass only)? How about other kernels?
Compare the runtime of the softmax operation versus the matrix multiplication operations within the self-attention layer of your model during a forward pass. How does the difference in runtimes compare to the difference in FLOPs?
Suppose we are training the model on a GPU and that the model parameters are originally in FP32. We’d like to use autocasting mixed precision with FP16. What are the data types of:(具体项目参见下方回答)
the model parameters within the autocast context? 全都是 torch.float32
the output of the first feed-forward layer (ToyModel.fc1)? torch.float16
the output of layer norm (ToyModel.ln)? torch.float16
the model’s predicted logits? torch.float16
the loss? torch.float16
the model’s gradients? 全都是 torch.float32
You should have seen that FP16 mixed precision autocasting treats the layer normalization layer differently than the feed-forward layers. What parts of layer normalization are sensitive to mixed precision? If we use BF16 instead of FP16, do we still need to treat layer normalization differently? Why or why not?
Modify your benchmarking script to optionally run the model using mixed precision with BF16. Time the forward and backward passes with and without mixed-precision for each language model size described in Section 2.1.2. Compare the results of using full precision versus mixed precision, and comment on any trends as model size changes. You may find the nullcontext no-op context manager to be useful.
Add an option to your profiling script to run your model through the memory profiler.
It may be helpful to reuse some of your previous infrastructure (e.g., to activate mixed-precision, load specific model sizes, etc). Then, run your script to get a memory profile of the xl model when either doing inference only (just forward pass) or a full training step. What do your memory timelines look like? Can you tell which stage is running based on the peaks you see?
上图是训练全流程的大致图像长这样(显存不够,使用 medium 代替,后文同理),最大的峰值应该是 backward 开始了一段时间的时候,GPU 需要绝大多数的激活值来算 grad,还需要申请新 grad 的空间。
What is the peak memory usage of each context length when doing a forward pass? What about when doing a full training step?
见上文
Find the peak memory usage of the xl model when using mixed-precision, for both a forward pass and a full training step. Does mixed-precision significantly affect memory usage?
Consider the xl model. Given our reference hyperparameters, what is the size of a tensor of activations in the Transformer residual stream, in single-precision? Give this size in MiB (i.e., divide the number of bytes by 10242).
Now look closely at the “Active Memory Timeline” from pytorch.org/memory_viz of a memory snapshot of the xl model doing a forward pass. When you reduce the “Detail” level, the tool hides the smallest allocations to the corresponding level (e.g., putting “Detail” at 10% only shows the 10% largest allocations). What is the size of the largest allocations shown? Looking through the stack trace, can you tell where those allocations come from?
Consider a Transformer with N identical blocks stacked sequentially. Without any checkpointing, all N blocks’ worth of residuals are kept alive simultaneously, giving O(N) peak activation memory. We have a free hand to wrap any subset of the forward pass in checkpoint, including nesting checkpoint calls inside one another.
What checkpointing strategy minimizes peak activation memory, ignoring the compute cost? Describe how you would arrange the checkpoint calls (a code sketch is fine), and give the asymptotic peak activation memory and compute of your strategy as a function of N. Assume the residuals saved by a single block dominate any per-checkpoint bookkeeping.
从复杂度的角度来算,如果不考虑计算量,最节约的是多重递归直到最小的单位。例如将 N 个 block 设置一个 ckpt,然后该块中进一步设置两个 ckpt,一直递归下去。
Consider the xl model config with batch size 4 and sequence length 2048 as above. If you only have the time/compute budget to run one step of recomputation (meaning you may not nest checkpoint calls), what is the best checkpointing strategy to reduce peak memory? Profile your run’s peak memory to validate your hypothesis. Compare the peak memory of the next smaller and larger checkpointing block sizes to be sure.
Benchmark your attention implementation at different scales. Write a script that will:
Fix the batch size to 8 and don’t use multihead attention (i.e. remove the head dimension).
Iterate through the cartesian product of [16, 32, 64, 128] for the head embedding dimension d model, and [256, 1024, 4096, 8192, 16384] for the sequence length.
Create random inputs Q,K,V for the appropriate size.
Time 100 forward passes through attention using the inputs.
Measure how much memory is in use before the backward pass starts, and time 100 backward passes.
Make sure to warm up, and to call torch.cuda.synchronize() after each forward/backward pass.
Depending on your GPU, some of these configurations are expected to run out of memory. Report the timings (or out-of-memory errors) you get for these configurations. At what size do you get out-of-memory errors? Do the accounting for the memory usage of attention in one of the smallest configurations you find that runs out of memory (you can use the equations for memory usage of Transformers from Assignment 1). How does the memory saved for backward change with the sequence length? What would you do to eliminate this memory cost?
Extend your attention benchmarking script to include a compiled version of your PyTorch implementation of attention, and compare its performance to the uncompiled version with the same configuration as the pytorch_attention problem above.
Now, compile your entire Transformer model in your end-to-end benchmarking script. How does the performance of the forward pass change? What about the combined forward and backward passes and optimizer steps?