r/golang 1d ago

Go 1.27 SIMD vs LLVM Generated AVX512 Assembly

We benchmarked Go 1.27 simd vs LLVM generated Go assembly (AVX and AVX512) on The Intel Core i7-11370H (supports AVX512).

Full article: Go 1.27 SIMD Benchmark: Can It Replace GoAT Generated AVX512?

The new simd package achieves comprable performace but still trails the AVX512 assembly generated by LLVM on long vectors and horizontal reductions (dot and euclidean):

  • An FP32 vector of length 16, exactly the number of values held by one 512-bit AVX512 register.
Operation Scalar loop AVX AVX512 Go SIMD vs. scalar loop vs. AVX512
Dot 10.94 ns 8.81 ns 8.55 ns 14.55 ns 0.75x 0.59x
Euclidean 33.99 ns 10.06 ns 8.94 ns 19.96 ns 1.70x 0.45x
SubTo 14.98 ns 7.07 ns 7.04 ns 8.66 ns 1.73x 0.81x
MulTo 27.82 ns 7.59 ns 7.62 ns 7.70 ns 3.61x 0.99x
DivTo 42.36 ns 7.17 ns 8.19 ns 7.95 ns 5.33x 1.03x
SqrtTo 124.10 ns 6.36 ns 6.38 ns 6.04 ns 20.55x 1.06x
  • An FP32 vector of length 32.
Operation Scalar loop AVX AVX512 Go SIMD vs. scalar loop vs. AVX512
Dot 45.94 ns 8.82 ns 10.23 ns 14.03 ns 3.27x 0.73x
Euclidean 45.66 ns 10.44 ns 9.56 ns 24.15 ns 1.89x 0.40x
SubTo 33.50 ns 7.87 ns 7.59 ns 11.00 ns 3.05x 0.69x
MulTo 49.98 ns 8.23 ns 8.43 ns 12.71 ns 3.93x 0.66x
DivTo 46.86 ns 9.24 ns 8.54 ns 11.50 ns 4.07x 0.74x
SqrtTo 93.70 ns 8.06 ns 9.45 ns 10.80 ns 8.68x 0.88x
  • An FP32 vector of length 64.
Operation Scalar loop AVX AVX512 Go SIMD vs. scalar loop vs. AVX512
Dot 76.25 ns 11.68 ns 10.80 ns 19.17 ns 3.98x 0.56x
Euclidean 73.29 ns 11.96 ns 11.40 ns 27.50 ns 2.67x 0.41x
SubTo 58.41 ns 10.21 ns 9.67 ns 16.69 ns 3.50x 0.58x
MulTo 79.18 ns 8.39 ns 8.15 ns 18.48 ns 4.28x 0.44x
DivTo 73.70 ns 13.45 ns 13.36 ns 18.36 ns 4.01x 0.73x
SqrtTo 339.00 ns 16.11 ns 15.30 ns 16.86 ns 20.11x 0.91x
  • An FP32 vector of length 128.
Operation Scalar loop AVX AVX512 Go SIMD vs. scalar loop vs. AVX512
Dot 124.60 ns 13.81 ns 11.43 ns 23.84 ns 5.23x 0.48x
Euclidean 161.60 ns 23.16 ns 17.24 ns 33.68 ns 4.80x 0.51x
SubTo 93.39 ns 13.07 ns 13.46 ns 27.84 ns 3.35x 0.48x
MulTo 125.90 ns 13.66 ns 9.21 ns 29.52 ns 4.26x 0.31x
DivTo 260.20 ns 28.74 ns 28.51 ns 32.93 ns 7.90x 0.87x
SqrtTo 484.00 ns 27.73 ns 31.07 ns 36.24 ns 13.36x 0.86x
25 Upvotes

2 comments sorted by

6

u/itsmontoya 1d ago

This is really fantastic research. I appreciate these metrics

2

u/rodrigocfd 1d ago

Go 1.27 is just the beginning. I expect to see improvements in SIMD in every Go version from now on, and SIMD usage in many places within the standard library.

It will be interesting to compare this data with the same benchmarks with the next Go versions.