PyTorch / Inductor
Add global-scratch support to PyTorch's static Triton launcher
Device-side TMA kernels could not launch from cached FX graphs because the static Triton launcher rejected global-scratch metadata. The landed change preserves scratch size and alignment, allocates workspace for the full launch grid, passes it through the kernel ABI, and keeps scratch kernels off the incompatible fast launcher.