FP8 is ~100 tflops faster when the kernel name has "cutlass" in it

(twitter.com)

172 points | by limoce 6 hours ago ago

75 comments