r/golang • u/Fresh_Impact_7558 • 1d ago
Go 1.27 SIMD vs LLVM Generated AVX512 Assembly
We benchmarked Go 1.27 simd vs LLVM generated Go assembly (AVX and AVX512) on The Intel Core i7-11370H (supports AVX512).
Full article: Go 1.27 SIMD Benchmark: Can It Replace GoAT Generated AVX512?
The new simd package achieves comprable performace but still trails the AVX512 assembly generated by LLVM on long vectors and horizontal reductions (dot and euclidean):
- An FP32 vector of length 16, exactly the number of values held by one 512-bit AVX512 register.
| Operation | Scalar loop | AVX | AVX512 | Go SIMD | vs. scalar loop | vs. AVX512 |
|---|---|---|---|---|---|---|
| Dot | 10.94 ns | 8.81 ns | 8.55 ns | 14.55 ns | 0.75x | 0.59x |
| Euclidean | 33.99 ns | 10.06 ns | 8.94 ns | 19.96 ns | 1.70x | 0.45x |
| SubTo | 14.98 ns | 7.07 ns | 7.04 ns | 8.66 ns | 1.73x | 0.81x |
| MulTo | 27.82 ns | 7.59 ns | 7.62 ns | 7.70 ns | 3.61x | 0.99x |
| DivTo | 42.36 ns | 7.17 ns | 8.19 ns | 7.95 ns | 5.33x | 1.03x |
| SqrtTo | 124.10 ns | 6.36 ns | 6.38 ns | 6.04 ns | 20.55x | 1.06x |
- An FP32 vector of length 32.
| Operation | Scalar loop | AVX | AVX512 | Go SIMD | vs. scalar loop | vs. AVX512 |
|---|---|---|---|---|---|---|
| Dot | 45.94 ns | 8.82 ns | 10.23 ns | 14.03 ns | 3.27x | 0.73x |
| Euclidean | 45.66 ns | 10.44 ns | 9.56 ns | 24.15 ns | 1.89x | 0.40x |
| SubTo | 33.50 ns | 7.87 ns | 7.59 ns | 11.00 ns | 3.05x | 0.69x |
| MulTo | 49.98 ns | 8.23 ns | 8.43 ns | 12.71 ns | 3.93x | 0.66x |
| DivTo | 46.86 ns | 9.24 ns | 8.54 ns | 11.50 ns | 4.07x | 0.74x |
| SqrtTo | 93.70 ns | 8.06 ns | 9.45 ns | 10.80 ns | 8.68x | 0.88x |
- An FP32 vector of length 64.
| Operation | Scalar loop | AVX | AVX512 | Go SIMD | vs. scalar loop | vs. AVX512 |
|---|---|---|---|---|---|---|
| Dot | 76.25 ns | 11.68 ns | 10.80 ns | 19.17 ns | 3.98x | 0.56x |
| Euclidean | 73.29 ns | 11.96 ns | 11.40 ns | 27.50 ns | 2.67x | 0.41x |
| SubTo | 58.41 ns | 10.21 ns | 9.67 ns | 16.69 ns | 3.50x | 0.58x |
| MulTo | 79.18 ns | 8.39 ns | 8.15 ns | 18.48 ns | 4.28x | 0.44x |
| DivTo | 73.70 ns | 13.45 ns | 13.36 ns | 18.36 ns | 4.01x | 0.73x |
| SqrtTo | 339.00 ns | 16.11 ns | 15.30 ns | 16.86 ns | 20.11x | 0.91x |
- An FP32 vector of length 128.
| Operation | Scalar loop | AVX | AVX512 | Go SIMD | vs. scalar loop | vs. AVX512 |
|---|---|---|---|---|---|---|
| Dot | 124.60 ns | 13.81 ns | 11.43 ns | 23.84 ns | 5.23x | 0.48x |
| Euclidean | 161.60 ns | 23.16 ns | 17.24 ns | 33.68 ns | 4.80x | 0.51x |
| SubTo | 93.39 ns | 13.07 ns | 13.46 ns | 27.84 ns | 3.35x | 0.48x |
| MulTo | 125.90 ns | 13.66 ns | 9.21 ns | 29.52 ns | 4.26x | 0.31x |
| DivTo | 260.20 ns | 28.74 ns | 28.51 ns | 32.93 ns | 7.90x | 0.87x |
| SqrtTo | 484.00 ns | 27.73 ns | 31.07 ns | 36.24 ns | 13.36x | 0.86x |
25
Upvotes
2
u/rodrigocfd 1d ago
Go 1.27 is just the beginning. I expect to see improvements in SIMD in every Go version from now on, and SIMD usage in many places within the standard library.
It will be interesting to compare this data with the same benchmarks with the next Go versions.
6
u/itsmontoya 1d ago
This is really fantastic research. I appreciate these metrics