MoonMath challenge
Hopper tensor cores
Explore an efficient W4A16 linear layer on NVIDIA Hopper tensor cores.
NVIDIA Hopper tensor cores use FP16 multiply-accumulate operations and support FP8 operations.
Explain how one might use this to implement a W4A16 linear layer. Compare the throughput and memory gain with full FP16 and W8A16 linear layers, and identify the problems or costs of the approach.
