PyTorch on Apple Silicon: I Got 3x Faster Inference with Metal Backend (No CUDA Required) on December 25, 2025 apple silicon ml labthecode pytorch metal torch compile +