Define the representative training workload
Record the actual model architecture, input shapes or sequence lengths, batch semantics, optimizer and quality target. Identify one device’s resident tensors under any sharding or offload arrangement. Use matching layer formulas as checks against the framework’s unique parameter inventory, not as a universal model-size formula.